Video call management method and device, electronic equipment, storage medium and computer product

By performing category detection and network slicing switching of audio and video data of video calls in 5G network, the problems of low data transmission efficiency and large latency in video calls are solved, and more efficient data transmission and lower latency are achieved.

CN120166186APending Publication Date: 2025-06-17CHINA MOBILE M2M +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364426.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In 5G networks, the prior art transmits all audio and video data during video calls, resulting in a decrease in data transmission efficiency, thereby increasing delay.

Method used

When the user terminal receives the audio and video data, the video call category detection is performed, the target network slice is determined, and the slice switching command is sent to the user terminal, so that it can perform appropriate audio, video or audio and video data transmission during the video call.

Benefits of technology

By detecting the categories of video calls and switching network slices, data transmission efficiency is improved, delays during video calls are reduced, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166186A_ABST
    Figure CN120166186A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication, and provides a video call management method and device, electronic equipment, a storage medium and a computer product, and the method comprises the steps: carrying out the type detection of a video call based on audio and video data transmitted by a user terminal in a video call process; the audio and video data are sent by the user terminal through a first network slice in a video call process; the first network slice is a network slice used for audio and video data transmission; determining a target network slice from the first network slice, the second network slice and the third network slice based on the category detection result; the second network slice is a network slice used for audio data transmission; the third network slice is a network slice used for video data transmission; sending a slice switching instruction to the user terminal based on the target network slice; and the user terminal performs audio data, video data or audio and video data transmission in the video call process based on the slice switching instruction. According to the invention, the delay during the video call can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a video call management method, apparatus, electronic device, storage medium, and computer product. Background Art

[0002] Real-time audio and video communication enables users to achieve real-time voice and video communication regardless of their location, greatly improving the communication efficiency and communication methods of users. Real-time audio and video communication has become the most important tool in communication. With the rapid development of the fifth-generation mobile communication (5th Generation Mobile Networks, 5G) technology, the characteristics of low latency and high bandwidth have brought revolutionary changes to video calls. However, how to achieve more stable and low-latency video calls in the 5G network is still an important challenge faced by the current technology field.

[0003] In traditional technologies, video calls need to compress and encode the audio and video data collected by user terminals such as mobile phones and tablets for sending to the other party through the network, and then decode the data after the other party receives it. Although the current encoding technology for audio and video data can compress the data to a certain extent and reduce the required bandwidth, when transmitting the audio and video data, all the audio and video data is transmitted in full, resulting in a certain reduction in the data transmission efficiency, and thus there may be delays when making video calls. Summary of the Invention

[0004] This application aims to at least solve one of the technical problems existing in the related technologies. For this purpose, this application provides a video call management method, apparatus, electronic device, storage medium, and computer product, to solve the problem that when transmitting audio and video data currently, all the audio and video data is transmitted in full, resulting in a certain reduction in the data transmission efficiency, and to reduce the delay when making video calls.

[0005] According to the video call management method of the first aspect embodiment of this application, it includes: Receiving the audio and video data sent by the user terminal during a video call, and performing category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for transmitting audio and video data; Based on the category detection result of the video call, determining a target network slice from the first network slice, the second network slice, and the third network slice; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data; Send a slice switching instruction to the user terminal based on the target network slice; the user terminal transmits audio data, video data, or audio and video data during a video call based on the slice switching instruction.

[0006] According to an embodiment of the present application, the category detection of the video call based on the audio and video data includes: Perform audio detection based on the audio and video data to obtain a first result of the audio detection; Perform video scene change detection based on the audio and video data to obtain a second result of the video scene change detection; Determine the category detection result of the video call based on the first result and the second result.

[0007] According to an embodiment of the present application, the performing audio detection based on the audio and video data to obtain a first result of the audio detection includes: Determine the short-time energy of each frame of audio signal in the audio and video data; Determine the audio signal with short-time energy greater than a preset short-time energy threshold as an audible audio signal, and determine the audio signal with short-time energy less than or equal to the preset short-time energy threshold as a silent audio signal; If it is determined that the video call is silent within a continuous preset time based on each frame of audio signal, output a silent result; otherwise, output an audible result.

[0008] According to an embodiment of the present application, the performing video scene change detection based on the audio and video data to obtain a second result of the video scene change detection includes: If it is determined based on the video frames in the audio and video data that the video scene does not change within a continuous preset time of the video call, output a no-change result; otherwise, output a change result; wherein, whether the video scene changes is determined by comparing the difference value between two video frames with a preset difference threshold.

[0009] According to an embodiment of the present application, the determining the category detection result of the video call based on the first result and the second result includes: If the first result is a silent result and the second result is a no-change result, determine that the video call is silent and there is no video scene change; If the first result is an audible result and the second result is a change result, determine that the video call is audible and there is a video scene change; If the first result is a silent result and the second result is a change result, determine that the video call is silent but there is a video scene change; If the first result is an audible result and the second result is an unchanged result, it is determined that the video call is audible but there is no change in the video scene.

[0010] According to an embodiment of the present application, determining a target network slice from a first network slice, a second network slice, and a third network slice based on the category detection result of the video call includes: If the video call is silent and there is no change in the video scene, or the video call is audible and there is a change in the video scene, the first network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice; If the video call is audible but there is no change in the video scene, the second network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice; If the video call is silent but there is a change in the video scene, the third network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice.

[0011] The video call management device according to the embodiment of the second aspect of the present application includes: A detection module, configured to receive audio - video data sent by a user terminal during a video call, and perform category detection of the video call based on the audio - video data; the audio - video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for transmitting audio - video data; A determination module, configured to determine a target network slice from a first network slice, a second network slice, and a third network slice based on the category detection result of the video call; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data; A switching module, configured to send a slice switching instruction to the user terminal based on the target network slice; the user terminal performs audio data, video data, or audio - video data transmission during the video call based on the slice switching instruction.

[0012] The electronic device according to the embodiment of the third aspect of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the video call management method as described in any one of the above.

[0013] The storage medium according to the embodiment of the fourth aspect of the present application is a non - transient computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the video call management method as described in any one of the above.

[0014] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program which, when executed by a processor, implements the video call management method as described in any one of the above.

[0015] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects: When receiving audio and video data sent by a user terminal during a video call, by detecting the category of the video call based on the audio and video data, and based on the detection result of the video call category, determining a target network slice from a first network slice for audio and video data transmission, a second network slice for audio data transmission, and a third network slice for video data transmission, and then sending a slice switching instruction to the user terminal based on the target network slice, so that the user terminal can transmit audio data, video data, or audio and video data during the video call based on the slice switching instruction. By detecting the category of the video call, scene analysis of the video call can be realized, and then according to different call scenarios, the user terminal can be instructed to use a suitable network slice for video call data transmission, which can improve data transmission efficiency and reduce latency during video calls.

[0016] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic flowchart of the video call management method provided by the embodiments of the present application.

[0019] Figure 2 is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following further describes in detail the embodiments of the present application in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0021] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the embodiments of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0022] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.

[0023] In the embodiments of the present application, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "under" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0024] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0025] The present application provides a video call management method, device, electronic device, storage medium, and computer product.

[0026] Figure 1 It is a schematic flowchart of the video call management method provided by an embodiment of this application. As Figure 1 shown, the video call management method includes: Step 110: Receive the audio and video data sent by the user terminal during the video call, and perform category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through the first network slice during the video call; the first network slice is a network slice for transmitting audio and video data.

[0027] Step 120: Determine the target network slice from the first network slice, the second network slice, and the third network slice based on the category detection result of the video call; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data.

[0028] Step 130: Send a slice switching instruction to the user terminal based on the target network slice; the user terminal performs the transmission of audio data, video data, or audio and video data during the video call based on the slice switching instruction.

[0029] It should be noted that the video call management method provided by the embodiment of this application can be applied to the scenario of video calls. The execution subject of the video call management method provided by the embodiment of this application can be a computer device deployed with an edge computing node. The computer device can be, for example, a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It should be noted that all the data that needs to be obtained in this application is obtained through regular channels after being authorized by the relevant users.

[0030] A video call management device can be set or connected in the computer device of this application, so as to control the video call management device to execute the video call management method of this application.

[0031] Specifically, three network slices can be established in advance in this application, namely, the "5G network slice dedicated to audio communication", the "5G network slice dedicated to video communication", and the "5G network slice dedicated to audio and video communication". Among them, the 5G network slice dedicated to audio and video communication can be defined as the first network slice in this application, the 5G network slice dedicated to audio communication can be defined as the second network slice in this application, and the 5G network slice dedicated to video communication can be defined as the third network slice in this application.

[0032] In daily life, when the person making a video call puts a user terminal such as a mobile phone or tablet on the desktop, and the caller does other things at the same time (such as cooking, cleaning, etc.), so that he does not appear in the video screen, in this case, there will be a situation where the video scene does not change, but the audio exists. The dedicated 5G network slice for audio communication in this application is mainly aimed at this situation. In this case, only low-latency, high-definition audio data transmission is required, and users need stable and clear audio transmission quality. The creation process of the dedicated 5G network slice for audio communication is as follows: First, network planning is carried out, network topology is designed, and a unique Internet Protocol (IP) address segment is allocated to the dedicated 5G network slice for audio communication. Next, configure the dedicated 5G network slice for audio communication. Use the Network Slice Management Function (NSMF) tool to create a new slice instance and set the parameters of the slice as follows: Slice name: 5G network slice dedicated to audio communication; Slice identifier: Slice_ID_001; Bandwidth: 2 megabits per second (Mbps); Delay: 30 milliseconds (ms); Jitter: 2ms; Packet loss rate: less than 0.1%; In the dedicated 5G network slice instance for audio communication, Advanced Audio Coding (AAC) technology is used, and audio encoders and decoders are configured to provide high-quality audio transmission. The Real-time Transport Protocol (RTP) is used, and the corresponding port number is set to ensure the real-time and reliability of audio data. The jitter buffer and packet loss rate detection mechanism are introduced to improve the stability of audio transmission. Finally, the dedicated 5G network slice for audio communication is deployed on the network topology.

[0033] In daily life, when a person on a video call is making a video call, the video screen is changing but no sound is emitted. For example, when making a video call at an exhibition, the person at the exhibition just wants to transmit the screen to the receiver in real time, but does not speak in order to keep the exhibition quiet. Or when visiting or inspecting certain places, only the scene is changing, but there is no audio around. In this case, the video scene changes but there is no audio. The dedicated 5G network slice for video communication is mainly aimed at this situation. In this case, only high-definition video screens need to be transmitted, and audio does not need to be transmitted. The creation process of the dedicated 5G network slice for video communication is as follows: First, configure the dedicated 5G network slice for video communication. Use the NSMF tool to create a new slice instance and set the parameters of the slice as follows: Slice name: 5G network slice dedicated to video communication; Slice identifier: Slice_ID_002; Bandwidth: 50Mbps; Delay: less than 50 milliseconds; Jitter: less than 10 milliseconds; Packet loss rate: less than 0.1%; In the dedicated 5G network slice instance for video communication, deploy a video encoder and decoder that supports 1080P resolution and 30 frames per second, and configure the parameters of the encoder and decoder to support video transmission with 1080P resolution, 30 frames per second, and 5Mbps bit rate; Configure the RTP / Real-time Transport Control Protocol (RTCP) protocol for video transmission and set the corresponding port number. Since no audio is transmitted, you can only configure the protocols and ports related to video transmission. Finally, the dedicated 5G network slice for video communication is deployed on the network topology.

[0034] In daily life, when a video caller is making a video call, the video screen is changing, accompanied by the voice of the caller or the surrounding environment. In this case, the video scene will change and there will be audio. The dedicated 5G network slice for video communication is mainly aimed at this situation. In this case, the video call needs to transmit high-quality audio and video. The creation process of the dedicated 5G network slice for audio and video communication is as follows: First, configure the dedicated 5G network slice for audio and video communication. Use the NSMF tool to create a new slice instance and set the parameters of the slice as follows: Slice name: 5G network slice dedicated to audio and video communication; Slice identifier: Slice_ID_003; Bandwidth: 20 Mbps of bandwidth is allocated to the dedicated 5G network slice for audio and video communications (8 Mbps for video and 320 kilobits per second (kbps) for audio, with some bandwidth reserved for redundancy and expansion); Delay: less than 40 milliseconds; Jitter: less than 5 milliseconds; Packet loss rate: less than 0.1%; In a dedicated 5G network slice instance for audio - video communication, deploy high - performance audio - video encoders and decoders, which support the encoding and decoding of 1080P video and 320kbps audio. Configure the parameters of the encoders and decoders to support the transmission of video frames with a resolution of 1080P, 30 frames per second, and a bitrate of 5Mbps. Configure the RTP / RTCP protocol for audio - video transmission and set the corresponding port numbers. Finally, deploy the dedicated 5G network slice for audio - video communication on the network topology.

[0035] Furthermore, in this application, edge computing nodes can be respectively embedded into the "dedicated 5G network slice for audio communication", the "dedicated 5G network slice for video communication", and the "dedicated 5G network slice for audio - video communication".

[0036] Specifically, in this application, edge computing nodes can be deployed at the edge locations of the network (such as base stations, access points, data centers, etc.). Edge computing nodes should have functions such as computing, storage, and network services, and be able to support the processing requirements of specific services. By respectively embedding edge computing nodes into the "dedicated 5G network slice for audio communication", the "dedicated 5G network slice for video communication", and the "dedicated 5G network slice for audio - video communication", the network slices can utilize the computing power and storage resources of the edge computing nodes. Configure the interfaces and protocols between the network slices and the edge computing nodes to ensure that the network slices can efficiently access and utilize the resources of the edge computing nodes.

[0037] Specifically, the RTP real - time transmission protocol can be configured for audio - video transmission between the "dedicated 5G network slice for audio communication" and the edge computing nodes.

[0038] Configure the RTP / RTCP protocol for audio - video transmission between the "dedicated 5G network slice for video communication" and the edge computing nodes.

[0039] Configure the RTP / RTCP protocol for audio - video transmission between the "dedicated 5G network slice for audio - video communication" and the edge computing nodes.

[0040] In addition, configure the interfaces and protocols between the slices and the edge computing nodes to ensure that the slices can efficiently access and utilize the resources of the edge computing nodes.

[0041] Thus, the user terminal of the sender can perform data transmission with the user terminal of the receiver through the dedicated 5G network slice for audio communication, the dedicated 5G network slice for video communication, and the dedicated 5G network slice for audio - video communication, while the edge computing nodes can perform data transmission with the user terminal through the dedicated 5G network slice for audio - video communication.

[0042] By utilizing 5G network slicing, a dedicated network slice is provided for the audio and video communication system to ensure the communication bandwidth and quality.

[0043] Based on this, when a user uses their user terminal to initiate a video call to another user's user terminal, both user terminals initiating the video call will send a "dedicated 5G network slice for audio and video communication" selection request to the base station. After receiving the selection request, the base station selects the corresponding slice for access according to the slice name in the selection request. Thus, a video call can be established.

[0044] Furthermore, during the video call, the user terminal of the party that needs to send audio and video data, as the sender, can convert the audio signal collected by the microphone into an analog electrical signal, and convert the analog audio signal into a digital audio signal through an analog-to-digital converter (ADC). During this process, the audio signal is sampled and quantized to form a series of discrete values. Use an audio coding algorithm (such as AAC, Opus, etc.) to encode the digital audio signal. Among them, Opus is an efficient lossy audio coding algorithm designed to provide low-latency, high-quality audio transmission and storage solutions.

[0045] And, the video signal collected by the camera can be converted into an analog electrical signal; the analog video signal is converted into a digital video signal through an analog-to-digital converter (ADC). During this process, the video signal is sampled and quantized to form a series of discrete values. Use a video coding algorithm (such as H.264, H.265, etc.) to encode the digital video signal. During the encoding process, the video signal of the video call is divided into a series of video frames, and each frame contains image information within a certain period of time; the encoded video frames will be further compressed to reduce the data volume.

[0046] Further, after the video call starts, the sender can first use the dedicated 5G network slice for audio and video communication accessed to simultaneously send the encoded audio and video data to the user terminal of the receiver and the edge computing node. The user terminal of the receiver decodes the audio and video data to restore it to the original audio and video signal and plays it on the user terminal of the receiver.

[0047] After receiving the audio and video data, the edge computing node can detect the category of the video call through the audio and video data, and determine whether at least the corresponding sender needs to adjust the audio and video data transmission strategy according to the detected category.

[0048] Among them, the categories of video calls can include the following: Silent and no video scene change; With sound and video scene change; With sound but no video scene change; Silent but existing video scene changes.

[0049] Furthermore, the edge computing node can determine a suitable network slice as the target network slice from the first network slice, the second network slice, and the third network slice according to the detection result of the video call category.

[0050] Further, the edge computing node can return a slice switching instruction to the user terminal of the sender according to the determined target network slice, so that the user terminal of the sender can switch the audio and video transmission policy according to the slice switching instruction.

[0051] Among them, the three slice switching instructions are as follows: Slice switching instruction 1: Switch to the "5G network slice dedicated to audio and video communication".

[0052] Slice switching instruction 2: Switch to the "5G network slice dedicated to audio communication".

[0053] Slice switching instruction 3: Switch to the "5G network slice dedicated to video communication".

[0054] The user terminal of the sender adjusts the audio and video data transmission policy of the video call according to the received slice switching instruction. The specific policies are as follows: Policy 1: When receiving slice switching instruction 1, the user terminal of the sender uses the "5G network slice dedicated to audio and video communication" to send the encoded audio and video data to the user terminal of the receiver and the edge computing node at the same time. Among them, if the "5G network slice dedicated to audio and video communication" has been used for transmission currently, there is no need to switch.

[0055] Policy 2: When receiving slice switching instruction 2, the user terminal of the sender uses the "5G network slice dedicated to video communication" to send only the encoded audio data to the user terminal of the receiver, and at the same time, the user terminal of the sender uses the "5G network slice dedicated to audio and video communication" to send the encoded audio and video data to the edge computing node.

[0056] Policy 3: When receiving slice switching instruction 3, the user terminal of the sender uses the "5G network slice dedicated to audio communication" to send only the encoded video data to the user terminal of the receiver, and at the same time, the user terminal of the sender uses the "5G network slice dedicated to audio and video communication" to send the encoded audio and video data to the edge computing node.

[0057] In summary, Policy 1 is applicable to the situation where there are both scene changes and sounds during a video call. For example, when two people are talking face to face, after adopting Policy 1, the video and audio will be played synchronously on the user terminal of the receiver.

[0058] Strategy 2 applies to the situation where the video scenario remains unchanged and there is audio. For example, when the video call participants place the user terminal on the table and the caller goes to do other things (such as cooking, cleaning, etc.) and thus does not appear in the video frame. In this case, it is not necessary to transmit the audio and video. Therefore, the user terminal of the sender will only send audio data and will not send video data. The user terminal of the receiver will continuously play the previous frame of the picture, but there is audio. By reducing the amount of video data sent and through a dedicated 5G network slice for audio communication, it is possible to improve the data transmission efficiency, reduce network latency and jitter.

[0059] Strategy 3 applies to the situation where the video scenario changes and there is no audio. For example, when having a video call in a place like an art exhibition, the person at the art exhibition just wants to transmit the picture to the receiver in real time and will not speak in order to keep the art exhibition quiet. Or when visiting or patrolling certain places, only the scene is changing and there is no audio around. In this case, it is not necessary to transmit audio. Therefore, the user terminal of the sender will only send video data and will not send audio data. The user terminal of the receiver will play the picture but there is no audio. By reducing the amount of audio data sent and through a dedicated 5G network slice for video communication, it is possible to improve the data transmission efficiency, reduce network latency and jitter.

[0060] Since the edge computing node needs to continuously analyze the audio and video of the video call, it is necessary to use a "dedicated 5G network slice for audio and video communication" to send the encoded audio and video data to the edge computing node.

[0061] Through 5G network slicing technology, audio recognition technology, scene recognition technology and edge computing technology, it is possible to analyze the video call scenario, adopt a suitable 5G network slice to transmit data according to different situations, improve the data transmission efficiency, reduce network latency and jitter, and improve the user's video call experience.

[0062] According to the video call management method of the embodiments of the present application, when receiving the audio-visual data sent by the user terminal during a video call, by detecting the category of the video call based on the audio-visual data, and based on the detection result of the category of the video call, determine the target network slice from the first network slice for audio-visual data transmission, the second network slice for audio data transmission, and the third network slice for video data transmission, and then send a slice switching instruction to the user terminal based on the target network slice, so that the user terminal can transmit audio data, video data, or audio-visual data during the video call based on the slice switching instruction. By detecting the category of the video call, the scenario analysis of the video call can be realized, and then according to different call scenarios, the user terminal can be instructed to use a suitable network slice for data transmission of the video call, which can improve the data transmission efficiency and reduce the delay during the video call.

[0063] Based on the above embodiments, detecting the category of the video call based on the audio-visual data includes: Performing audio detection based on the audio-visual data to obtain a first result of the audio detection; Performing video scene change detection based on the audio-visual data to obtain a second result of the video scene change detection; Based on the first result and the second result, determine the detection result of the category of the video call.

[0064] Specifically, when the present application detects the category of the video call based on the audio-visual data, it can first perform audio detection according to the audio-visual data, and determine the obtained result as the first result. Specifically, it can be determined by comparing the short-time energy of each frame of audio signal in the audio-visual data with the short-time energy of other frames of audio signal in the current video call, and the comparison result with a preset short-time energy threshold T.

[0065] And, it can perform video scene change detection according to the audio-visual data, and determine the obtained result as the second result. Specifically, it can be determined by comparing the difference value between consecutive two video frames of each video frame in the audio-visual data and other video frames in the current video call with a preset difference threshold.

[0066] Furthermore, according to the first result of the audio detection and the second result of the video scene change detection, it can be determined whether the category of the video call is silent and there is no video scene change, there is sound and there is a video scene change, there is sound but there is no video scene change, or there is no sound but there is a video scene change. So as to facilitate subsequent determination of the target network slice from the first network slice, the second network slice, and the third network slice according to the detection result of the category of the video call, and then sending a slice switching instruction to the user terminal based on the target network slice.

[0067] By accurately determining the category of the video call, the present application can accurately return a slice switching instruction to the user terminal of the sender, enabling the user terminal of the sender to switch an appropriate audio-video transmission strategy according to the slice switching instruction, which can improve the data transmission efficiency and reduce the latency during the video call. Moreover, through edge computing acceleration processing, the real-time response ability can be improved.

[0068] Based on the above embodiments, audio detection is performed on the audio-video data to obtain a first result of the audio detection, including: Determine the short-time energy of each frame of audio signal in the audio-video data; Determine the audio signal with short-time energy greater than the preset short-time energy threshold as the audible audio signal, and determine the audio signal with short-time energy less than or equal to the preset short-time energy threshold as the silent audio signal; If it is determined that the video call is silent for a continuous preset time based on each frame of audio signal, output a silent result; otherwise, output an audible result.

[0069] Specifically, the edge computing node of the present application can decode the audio-video data transmitted by the user terminal of the sender, divide the audio signal therein into a series of equally long frames, and the length of each frame is usually the same as the length N of the window function. In this way, each frame of audio signal can be regarded as a short-time audio signal segment.

[0070] Furthermore, for each frame of audio signal, its short-time energy E(n) can be calculated. The short-time energy E(n) can be specifically calculated by the following formula: E(n)=Σ[x(m)*w(n-m)]², where m ranges from n - N + 1 to n; w(n) is a window function, and N is the length of the window function.

[0071] It should be noted that the edge computing node of the present application can also call other partial frame audio signals and their short-time energies before the current time as needed.

[0072] In the present application, a short-time energy threshold T can be preset in advance. When the short-time energy E(n) of a certain frame of audio signal is greater than the threshold T, it is considered that the frame of audio signal contains sound; otherwise, it is considered that the frame of audio signal is silent.

[0073] Traverse each frame of audio signal required, and judge whether each frame of audio signal contains sound according to the comparison result of its short-time energy and the threshold T. When the video call is silent for a continuous preset time (for example, 10 seconds), output a silent result. That is, the user of the user terminal of the sender does not make a sound for a long time.

[0074] If any frame of the audio signal contains sound within a continuous preset time (such as 10 seconds), an audible result can be output. That is to say, it indicates that the user has not been silent for a long time.

[0075] In this application, the edge computing node decodes the audio and video data transmitted by the sender's user terminal, divides the audio signal into a series of equally long frames, traverses all the audio frame signals, and determines whether each audio frame signal contains sound according to the comparison result between its short-time energy and the threshold value. It can determine whether someone is speaking during a video call, providing a basis for switching the 5G network slice. Furthermore, it can instruct the user terminal to use a suitable network slice for data transmission during a video call according to different call scenarios, which can improve the data transmission efficiency and reduce the latency during a video call.

[0076] Based on the above embodiments, video scene change detection is performed based on the audio and video data, and a second result of the video scene change detection is obtained, including: If it is determined based on the video frames in the audio and video data that the video scene has not changed within a continuous preset time during the video call, an unchanged result is output; otherwise, a changed result is output; wherein, whether the video scene has changed is determined by comparing the difference value between two video frames with a preset difference threshold. Specifically, after the edge computing node of this application decodes the audio and video data transmitted by the sender's user terminal, a video frame sequence can be obtained. And other partial video frames before the current time can be called as needed.

[0077] Furthermore, for two consecutive video frames, the difference value between the two video frames can be calculated, which can be specifically implemented by the following formula: D = |F1 - F2|; where D represents the difference value between two video frames F1 and F2, and | | represents taking the absolute value.

[0078] When the value of D exceeds the preset difference threshold, it can be considered that the scene of the video call has changed; otherwise, it is considered that the scene of the video call has not changed.

[0079] Furthermore, according to the comparison result between each difference value and the preset difference threshold, it can be determined whether the video scene has changed within a continuous preset time (such as 10 seconds) during the video call.

[0080] If it is determined that the video scene has not changed within a continuous preset time (such as 10 seconds) during the video call, it indicates that the user is not within the call screen and is likely not to appear in the screen in the next short period of time. Therefore, an unchanged result can be output.

[0081] If it is determined that the video scene has changed during a continuous preset time (e.g., 10 seconds) of the video call, it indicates that the user has not been away from the call screen for a long time. Therefore, a result of "changed" can be output.

[0082] In this application, the audio and video data transmitted by the sending user terminal is decoded at the edge computing node, and whether the scene has changed is accurately judged according to the decoded video sequence, providing a basis for switching the 5G network slice. Furthermore, according to different call scenarios, the user terminal can be instructed to use a suitable network slice for data transmission of the video call, which can improve the data transmission efficiency and reduce the delay during the video call.

[0083] Based on the above embodiments, based on the first result and the second result, determine the category detection result of the video call, including: If the first result is a silent result and the second result is a no-change result, it is determined that the video call is silent and there is no video scene change; If the first result is an audible result and the second result is a changed result, it is determined that the video call is audible and there is a video scene change; If the first result is a silent result and the second result is a changed result, it is determined that the video call is silent but there is a video scene change; If the first result is an audible result and the second result is a no-change result, it is determined that the video call is audible but there is no video scene change.

[0084] Specifically, if this application determines that the first result is a silent result and the second result is a no-change result, it indicates that the user is not in the call screen and has not made a sound nearby. Therefore, it can be determined that the video call is silent and there is no video scene change.

[0085] If it is determined that the first result is an audible result and the second result is a changed result, it indicates that the user is in the call screen and has made a sound. Therefore, it can be determined that the video call is audible and there is a video scene change.

[0086] If it is determined that the first result is a silent result while the second result is a changed result, it indicates that the user is in the screen but has not made a sound. Therefore, it can be determined that the video call is silent but there is a video scene change.

[0087] If it is determined that the first result is an audible result and the second result is a no-change result, it indicates that the user is not in the call screen but has made a sound. Therefore, it can be determined that the video call is audible but there is no video scene change.

[0088] This application accurately determines the category detection of a video call, enabling the determination of a target network slice from a first network slice for audio-video data transmission, a second network slice for audio data transmission, and a third network slice for video data transmission based on the category detection result of the video call. Furthermore, a slice switching instruction is sent to the user terminal based on the target network slice, allowing the user terminal to transmit audio data, video data, or audio-video data during the video call based on the slice switching instruction. By detecting the category of the video call, scenario analysis of the video call can be achieved, and then the user terminal can be instructed to use a suitable network slice for video call data transmission according to different call scenarios, which can improve the data transmission efficiency and reduce the latency during the video call.

[0089] Based on the above embodiments, determining a target network slice from the first network slice, the second network slice, and the third network slice based on the category detection result of the video call includes: If the video call is silent and there is no video scene change, or the video call has sound and there is a video scene change, the first network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice; If the video call has sound but there is no video scene change, the second network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice; If the video call is silent but there is a video scene change, the third network slice among the first network slice, the second network slice, and the third network slice is determined as the target network slice.

[0090] Specifically, if this application determines that the video call is silent and there is no video scene change, or determines that the video call has sound and there is a video scene change, the first network slice among the first network slice, the second network slice, and the third network slice can be determined as the target network slice.

[0091] If it is determined that the video call has sound but there is no video scene change, the second network slice among the first network slice, the second network slice, and the third network slice can be determined as the target network slice.

[0092] If it is determined that the video call is silent but there is a video scene change, the third network slice among the first network slice, the second network slice, and the third network slice can be determined as the target network slice.

[0093] This application accurately determines a target network slice by detecting whether there is sound in the audio and video of a video call and whether the scene has changed, generating three slice switching instructions, enabling the user terminal of the sender to adjust the audio and video data transmission strategy of the video call according to the received slice switching instruction, so as to avoid sending unnecessary data, thereby improving the data transmission efficiency, reducing network latency and jitter, and enhancing the user's video call experience.

[0094] The video call management device provided by this application will be described below. The video call management device described below can be correspondingly referred to the video call management method described above.

[0095] Furthermore, this application also provides a video call management device.

[0096] The video call management device includes: A detection module, configured to receive audio and video data sent by a user terminal during a video call, and perform category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for audio and video data transmission; A determination module, configured to determine a target network slice from a first network slice, a second network slice, and a third network slice based on the category detection result of the video call; the second network slice is a network slice for audio data transmission; the third network slice is a network slice for video data transmission; A switching module, configured to send a slice switching instruction to the user terminal based on the target network slice; the user terminal performs audio data, video data, or audio and video data transmission during the video call based on the slice switching instruction.

[0097] When the video call management device of this application receives the audio and video data sent by the user terminal during the video call, it performs category detection of the video call based on the audio and video data, and determines a target network slice from the first network slice for audio and video data transmission, the second network slice for audio data transmission, and the third network slice for video data transmission based on the category detection result of the video call. Then, it sends a slice switching instruction to the user terminal based on the target network slice, enabling the user terminal to perform audio data, video data, or audio and video data transmission during the video call based on the slice switching instruction. By detecting the category of the video call, the scene analysis of the video call can be realized, and then the user terminal can be instructed to use a suitable network slice for video call data transmission according to different call scenarios, which can improve the data transmission efficiency and reduce the latency during the video call.

[0098] In one embodiment, the detection module is specifically configured to: Perform audio detection based on the audio-video data to obtain a first result of the audio detection; Perform video scene change detection based on the audio-video data to obtain a second result of the video scene change detection; Determine a category detection result of the video call based on the first result and the second result.

[0099] In one embodiment, the detection module is further configured to: Determine the short-time energy of each frame of audio signal in the audio-video data; Determine the audio signal with short-time energy greater than a preset short-time energy threshold as an audible audio signal, and determine the audio signal with short-time energy less than or equal to the preset short-time energy threshold as a silent audio signal; If it is determined that the video call is silent for a continuous preset time based on each frame of audio signal, output a silent result; otherwise, output an audible result.

[0100] In one embodiment, the detection module is further configured to: If it is determined that the video scene has not changed for a continuous preset time in the video call based on the video frames in the audio-video data, output a no-change result; otherwise, output a change result; wherein, whether the video scene has changed is determined by comparing the difference value between two video frames with a preset difference threshold.

[0101] In one embodiment, the detection module is further configured to: If the first result is a silent result and the second result is a no-change result, determine that the video call is silent and there is no video scene change; If the first result is an audible result and the second result is a change result, determine that the video call is audible and there is a video scene change; If the first result is a silent result and the second result is a change result, determine that the video call is silent but there is a video scene change; If the first result is an audible result and the second result is a no-change result, determine that the video call is audible but there is no video scene change.

[0102] In one embodiment, the determination module is specifically configured to: If the video call is silent and there is no video scene change, or the video call is audible and there is a video scene change, determine the first network slice among the first network slice, the second network slice, and the third network slice as the target network slice; If there is sound in the video call but no video scene change, determine the second network slice among the first network slice, the second network slice, and the third network slice as the target network slice; If there is no sound in the video call but there is a video scene change, determine the third network slice among the first network slice, the second network slice, and the third network slice as the target network slice.

[0103] Figure 2 An example of a schematic diagram of the physical structure of an electronic device is shown as Figure 2 shown. The electronic device may include: a processor 210, a communications interface 220, a memory 230, and a communication bus 240. Among them, the processor 210, the communications interface 220, and the memory 230 communicate with each other through the communication bus 240. The processor 210 may call the logical instructions in the memory 230 to execute the following method: receiving audio and video data sent by the user terminal during a video call, and performing category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for transmitting audio and video data; Based on the category detection result of the video call, determine the target network slice from the first network slice, the second network slice, and the third network slice; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data; Send a slice switching instruction to the user terminal based on the target network slice; the user terminal transmits audio data, video data, or audio and video data during the video call based on the slice switching instruction.

[0104] In addition, when the logical instructions in the above-mentioned memory 230 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application essentially, or the part that contributes to the related technology, or this part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0105] In another aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above embodiments. For example, it includes: receiving audio and video data sent by a user terminal during a video call, and performing category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for transmitting audio and video data; Based on the category detection result of the video call, determining a target network slice from the first network slice, the second network slice, and the third network slice; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data; Sending a slice switching instruction to the user terminal based on the target network slice; the user terminal performs audio data, video data, or audio and video data transmission during the video call based on the slice switching instruction.

[0106] In another aspect, an embodiment of the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above embodiments. For example, it includes: receiving audio and video data sent by a user terminal during a video call, and performing category detection of the video call based on the audio and video data; the audio and video data is sent by the user terminal through a first network slice during the video call; the first network slice is a network slice for transmitting audio and video data; Based on the category detection result of the video call, determining a target network slice from the first network slice, the second network slice, and the third network slice; the second network slice is a network slice for transmitting audio data; the third network slice is a network slice for transmitting video data; Sending a slice switching instruction to the user terminal based on the target network slice; the user terminal performs audio data, video data, or audio and video data transmission during the video call based on the slice switching instruction.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the related technologies can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application.

Claims

1. A video call management method, characterized in that: include: receiving audio and video data sent by a user terminal during a video call, and detecting a type of the video call based on the audio and video data; The audio and video data is sent by the user terminal through the first network slice during the video call process; The first network slice is a network slice used for audio and video data transmission; Based on the category detection result of the video call, determine the target network slice from the first network slice, the second network slice, and the third network slice; The second network slice is a network slice for transmitting audio data; The third network slice is a network slice for transmitting video data; A slice switching instruction is sent to the user terminal based on the target network slice; and the user terminal transmits audio data, video data, or audio and video data during a video call based on the slice switching instruction.

2. The video call management method according to claim 1, characterized in that: The detecting the type of the video call based on the audio and video data includes: Performing audio detection based on the audio and video data to obtain a first result of the audio detection; Performing video scene change detection based on the audio and video data to obtain a second result of the video scene change detection; A category detection result of the video call is determined based on the first result and the second result.

3. The video call management method according to claim 2, characterized in that: The performing audio detection based on the audio and video data to obtain a first result of the audio detection includes: Determining the short-time energy of each frame of audio signal in the audio and video data; Determine an audio signal with a short-time energy greater than a preset short-time energy threshold as a sound audio signal, and determine an audio signal with a short-time energy less than or equal to the preset short-time energy threshold as a silent audio signal; If it is determined based on each frame of audio signal that the video call is silent for a continuous preset time, a silent result is output; otherwise, an audible result is output.

4. The video call management method according to claim 3, characterized in that: The performing video scene change detection based on the audio and video data to obtain a second result of the video scene change detection includes: If it is determined based on the video frames in the audio and video data that the video scene of the video call has not changed within a continuous preset time, a no-change result is output; otherwise, a change result is output; wherein, whether the video scene has changed is determined by comparing the difference value between the two video frames with a preset difference threshold.

5. The video call management method according to claim 4, characterized in that: The determining the category detection result of the video call based on the first result and the second result includes: If the first result is a silent result and the second result is a no-change result, it is determined that the video call is silent and there is no video scene change; If the first result is a sound result and the second result is a change result, it is determined that the video call has sound and there is a video scene change; If the first result is a silent result and the second result is a changed result, it is determined that the video call is silent but there is a video scene change; If the first result is a sound result and the second result is a no-change result, it is determined that the video call has sound but there is no video scene change.

6. The video call management method according to claim 5, characterized in that: The determining, based on the category detection result of the video call, the target network slice from the first network slice, the second network slice, and the third network slice includes: If the video call is silent and there is no video scene change, or if the video call is in sound and there is a video scene change, determining the first network slice among the first network slice, the second network slice and the third network slice as the target network slice; If the video call has sound but there is no video scene change, the second network slice among the first network slice, the second network slice and the third network slice is determined as the target network slice; If the video call is silent but there is a video scene change, the third network slice among the first network slice, the second network slice and the third network slice is determined as the target network slice.

7. A video call management device, characterized in that: include: A detection module, configured to receive audio and video data sent by a user terminal during a video call, and perform video call category detection based on the audio and video data; The audio and video data is sent by the user terminal through the first network slice during the video call process; The first network slice is a network slice used for audio and video data transmission; A determination module, configured to determine a target network slice from the first network slice, the second network slice, and the third network slice based on a category detection result of the video call; The second network slice is a network slice for transmitting audio data; The third network slice is a network slice for transmitting video data; A switching module is used to send a slice switching instruction to the user terminal based on the target network slice; the user terminal transmits audio data, video data or audio and video data during the video call based on the slice switching instruction.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video call management method as described in any one of claims 1-6 is implemented.

9. A storage medium, the storage medium being a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, the video call management method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the video call management method according to any one of claims 1 to 6 is implemented.