Communication methods, devices and storage media

By acquiring and analyzing audio and video information from multi-party video calls, and prioritizing the display of important video footage, the issue of users needing to flip through pages during multi-party video calls is resolved, thus improving the user experience.

CN119316545BActive Publication Date: 2026-05-26HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2023-07-13
Publication Date
2026-05-26

Smart Images

  • Figure CN119316545B_ABST
    Figure CN119316545B_ABST
Patent Text Reader

Abstract

This application provides a call method, device, and storage medium. The method includes: acquiring first information of a second user corresponding to a second device, the first information including at least one of audio information or video information; acquiring second information of a third user corresponding to a third device, the second information including at least one of audio information or video information; determining the display positions of the second user's video frame and the third user's video frame on a multi-party video call interface in a first device based on the first information and the second information; and displaying the second user's video frame and the third user's video frame according to the determined display positions. The method provided by this application helps to improve the user's video call experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart terminals, and more particularly to a communication method, device, and storage medium. Background Technology

[0002] Multi-party calls refer to calls in which multiple users can participate in the same conversation. Typically, multi-party calls can include services such as video and audio.

[0003] Currently, in multi-party video call scenarios, when there are many participating users, multiple users may be displayed in separate pages, causing the main interface to only show the video feeds of a portion of the users. When a user is video calling with other participants, that user may need to scroll through pages to see the other participants' video feeds, thus degrading the user experience. Summary of the Invention

[0004] This application provides a calling method, device, and storage medium that help improve the user's video calling experience.

[0005] A first aspect provides a call method applied to a first device, wherein the first device conducts a multi-party video call with a second device and a third device, comprising: obtaining first information of a second user corresponding to the second device, the first information including at least one of audio information or video information; obtaining second information of a third user corresponding to the third device, the second information including at least one of audio information or video information; determining, based on the first information and the second information, the display positions of the video frames of the second user and the third user on the multi-party video call interface of the first device; and displaying the video frames of the second user and the third user according to the determined display positions.

[0006] In this application, by obtaining the audio and video information of the participants, the video footage of the participants is sorted and displayed on the homepage of the current user's device based on the audio and video information of the participants. This allows the video footage of the desired object to be displayed on the current user's homepage, thereby improving the user experience.

[0007] In one possible implementation, determining the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the first information and the second information includes: determining the priority of the second user's video feed and the priority of the third user's video feed based on the first information and the second information; and determining the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the priority of the second user's video feed and the priority of the third user's video feed.

[0008] In one possible implementation, determining the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface in the first device based on the priority of the second user's video feed and the third user's video feed includes: if the priority of the second user's video feed is higher than the priority of the third user's video feed, displaying the second user's video feed on the homepage of the multi-party video call interface in the first device, and displaying the third user's video feed on other pages of the multi-party video call interface; or, if the priority of the second user's video feed is higher than the priority of the third user's video feed, displaying the second user's video feed in a first area of ​​the homepage of the multi-party video call interface in the first device, and displaying the third user's video feed in a second area, wherein the homepage includes the first area and the second area, and the size of the first area is larger than the size of the second area.

[0009] In one possible implementation, the video information in the first information includes the posture of the second user, and the video information in the second information includes the posture of the third user; or, the audio information in the first information includes the volume of the second user, and the audio information in the second information includes the volume of the third user.

[0010] In one possible implementation, the volume of the second user is determined by multiple sample values ​​in the audio stream of the second user acquired by the second device, and the volume of the third user is determined by multiple sample values ​​in the audio stream of the third user acquired by the third device.

[0011] In one possible implementation, the audio information in the first information includes the speech content of the second user, and the audio information in the second information includes the speech content of the third user.

[0012] In one possible implementation, determining the priority of the second user's video feed and the priority of the third user's video feed based on the first information and the second information includes: determining whether the second user's speech content and the third user's speech content contain the first user corresponding to the first device; if the second user's speech content contains the first user and the third user's speech content does not contain the first user, the priority of the second user's video feed is higher than the priority of the third user's video feed; or, if the third user's speech content contains the first user and the second user's speech content does not contain the first user, the priority of the third user's video feed is higher than the priority of the second user's video feed.

[0013] In one possible implementation, determining the priority of the second user's video feed and the priority of the third user's video feed based on the first information and the second information includes: matching the speech content of the first user corresponding to the first device with the speech content of the second user and the speech content of the third user, respectively; if the matching degree between the speech content of the first user and the speech content of the second user is higher than the matching degree between the speech content of the first user and the speech content of the third user, the priority of the second user's video feed is higher than the priority of the third user's video feed; or, if the matching degree between the speech content of the first user and the speech content of the third user is higher than the matching degree between the speech content of the first user and the speech content of the second user, the priority of the third user's video feed is higher than the priority of the second user's video feed.

[0014] In one possible implementation, the method further includes: determining the target object in the speech content of the first user corresponding to the first device; if the target object in the speech content of the first user is the second user, the video frame of the second user has a higher priority than the video frame of the third user; or, if the target object in the speech content of the first user is the third user, the video frame of the third user has a higher priority than the video frame of the second user.

[0015] The second aspect provides a communication device including one or more functional modules for implementing the communication method as described in the first aspect.

[0016] A third aspect provides an apparatus comprising: a processor and a memory, the memory being used to store a program; the processor being used to run the program to implement the call method as described in the first aspect.

[0017] The fourth aspect provides a readable storage medium storing a program that, when run on a device, causes the device to implement the call method as described in the first aspect.

[0018] The fifth aspect provides a program that, when run on the processor of a device, causes the device to perform the call method as described in the first aspect.

[0019] In one possible design, the program in the fifth aspect can be stored wholly or partially on a storage medium packaged with the processor, or it can be stored wholly or partially on a memory not packaged with the processor. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the device provided in the embodiments of this application;

[0021] Figure 2 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;

[0022] Figure 3 A flowchart illustrating an embodiment of the call method provided in this application;

[0023] Figure 4a and Figure 4b This is a schematic diagram of a video display provided in an embodiment of this application;

[0024] Figure 5 A flowchart illustrating another embodiment of the call method provided in this application;

[0025] Figure 6 A schematic diagram of the structure of one embodiment of the communication device provided in this application. Detailed Implementation

[0026] In this embodiment of the application, unless otherwise stated, the character " / " indicates that the preceding and following objects are in an OR relationship. For example, A / B can represent A or B. "AND / OR" describes the relationship between the associated objects, indicating that three relationships can exist. For example, A AND / OR B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0027] It should be noted that the terms "first" and "second" used in the embodiments of this application are used only for distinguishing descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated, nor should they be construed as indicating or implying order.

[0028] In the embodiments of this application, "at least one" refers to one or more items, and "more than one" refers to two or more items. Furthermore, "at least one of the following" or similar expressions refer to any combination of these items, which may include any combination of a single item or a plurality of items. For example, at least one of A, B, or C can represent: A, B, C, A and B, A and C, B and C, or A, B, and C. Each of A, B, and C can be an element itself or a set containing one or more elements.

[0029] In this application, terms such as "exemplary," "in some embodiments," and "in another embodiment" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0030] In the embodiments of this application, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, their meanings are consistent. Similarly, in the embodiments of this application, "communication" and "transmission" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, their meanings are consistent. For example, transmission can include sending and / or receiving, and can be a noun or a verb.

[0031] In the embodiments of this application, the term "equal to" can be used in conjunction with "greater than" to apply to technical solutions employing the condition of "greater than", and can also be used in conjunction with "less than" to apply to technical solutions employing the condition of "less than". It should be noted that when "equal to" is used with "greater than", it cannot be used with "less than"; and when "equal to" is used with "less than", it cannot be used with "greater than".

[0032] Multi-party calls refer to calls in which multiple users can participate in the same conversation. Typically, multi-party calls can include services such as video and audio.

[0033] Currently, in multi-party video call scenarios, when there are many participating users, multiple users may be displayed in separate pages, causing the main interface to only show the video feeds of a portion of the users. When a user is video calling with other participants, that user may need to scroll through pages to see the other participants' video feeds, thus degrading the user experience.

[0034] Based on the above problems, this application proposes a call method that can be applied to a device to improve the user's video call experience.

[0035] The device includes a display screen, a camera, a speaker, and a microphone. The display screen is used to show video footage. The camera is used to capture user data from devices including, but not limited to: smart screens, desktop computers, laptops, mobile phones, tablets, vehicle-mounted terminals, user equipment (UE), access terminals, user units, user stations, mobile stations, mobile stations, remote stations, remote terminals, mobile devices, user terminals, terminals, wireless communication equipment, user agents, user devices, vehicle-mounted equipment, vehicle-to-everything (V2X) terminals, handheld communication devices, handheld computing devices, satellite wireless equipment, and other devices used for communication over wireless systems.

[0036] Figure 1 First, a schematic diagram of the structure of device 100 is shown as an example.

[0037] Device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0038] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on device 100. In other embodiments of this application, device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0039] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0040] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0041] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0042] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0043] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the device 100. In other embodiments of this application, the device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0044] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device via the power management module 141.

[0045] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0046] The wireless communication function of device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0047] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0048] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via the antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0049] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0050] The wireless communication module 160 can provide solutions for wireless communication applications on device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0051] In some embodiments, antenna 1 of device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling device 100 to communicate with networks and other devices via wireless communication technology.

[0052] Device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0053] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, device 100 may include one or N displays screens 194, where N is a positive integer greater than 1.

[0054] Device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0055] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye.

[0056] Camera 193 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Internet Service Provider) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0057] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0058] Video codecs are used to compress or decompress digital video. Device 100 may support one or more video codecs. Thus, device 100 can play or record video in various encoded formats.

[0059] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications such as image recognition, facial recognition, and speech recognition.

[0060] The external storage interface 120 can be used to connect an external storage card, such as a MicroSD card, to expand the storage capacity of the device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0061] Internal memory 121 can be used to store executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as a sound playback function), etc. The data storage area may store data created during the use of device 100 (such as audio data), etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device or flash memory device. Processor 110 executes various functional applications and data processing of device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor.

[0062] The device 100 can implement audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.

[0063] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0064] Speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Device 100 can listen to music or make hands-free calls through speaker 170A. Receiver 170B, also known as a "handpiece," is used to convert audio electrical signals into sound signals. When device 100 answers a phone call or voice message, the receiver 170B can be brought close to the user's ear to hear the voice. Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. Headphone jack 170D is used to connect wired headphones. Pressure sensor 180A is used to sense pressure signals and converts them into electrical signals. Gyroscope sensor 180B can be used to determine the motion posture of device 100. Barometric pressure sensor 180C is used to measure barometric pressure. Magnetometer 180D includes a Hall effect sensor. Accelerometer 180E can detect the magnitude of acceleration of device 100 in various directions (generally three axes). Distance sensor 180F is used to measure distance. Device 100 can measure distance using infrared or laser. In some embodiments, during a shooting scene, device 100 can utilize distance sensor 180F for distance measurement to achieve fast focusing. Proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photosensor, such as a photodiode. Ambient light sensor 180L is used to sense ambient light intensity. Fingerprint sensor 180H is used to collect fingerprints. Temperature sensor 180J is used to detect temperature. Touch sensor 180K, also called a "touch device," can be located on display screen 194, forming a touchscreen, also called a "touchscreen." Touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. Bone conduction sensor 180M can acquire vibration signals. Buttons 190 include a power button, volume buttons, etc. Motor 191 can generate vibration cues. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, or to indicate messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card.

[0065] Now combined Figure 2 Figure 4 illustrates the call method provided in the embodiments of this application.

[0066] Figure 2 This is an application scenario architecture diagram of the call method provided in an embodiment of this application. (See reference...) Figure 2 Application scenarios can include the first device, the second device, and the third device.

[0067] The number of second and third devices can be one or more, and the first, second, and third devices can form a multi-party video call scenario. In the multi-party video call scenario, a connection can be pre-established between the first device and the second and third devices. It is understood that the above connection can be made via a wired method or via a wireless method, and this application embodiment does not make any special limitation in this regard.

[0068] A connection can also be established between the second and third devices. Figure 2 (Not shown in the image), the connection method between the second and third devices can be referred to the connection method between the first and second devices and the third device, and will not be repeated here.

[0069] The first device can be a device used by a user, while the second and third devices can be devices used by other participating parties. The first, second, and third devices can be the aforementioned device 100.

[0070] It is understandable that the second and third devices are relative to the first device. That is, for any one of these devices, it can be considered the first device. From the perspective of the first device, the user using the first device can be the primary user, and users using other devices (e.g., the second and third devices) can be considered participants. In a multi-party video call scenario, each user can be the primary user. For ease of explanation, the user using the first device will be referred to as the "first user," the user using the second device as the "second user," and the user using the third device as the "third user."

[0071] During use, the first device can acquire information collected by the second and third devices, and can sort and display the video feeds of each participant based on this information. This allows the currently speaking participant to be displayed on the main interface, eliminating the need for users to manually flip through pages to view their video feed, thus improving the user's video call experience. The information collected by the second and third devices can include the participants' (e.g., the second and third users) posture information, volume levels, and the content of their speech.

[0072] The following text combines Figure 3 Figure 4 illustrates the call method provided in the embodiments of this application.

[0073] like Figure 3 The diagram shown is a flowchart of an embodiment of the call method provided in this application. Figure 3 In the illustrated embodiment, the first device can sort and display the video feeds of the participants based on the user posture information collected by the second and third devices, specifically including the following steps:

[0074] Step 301: The second device identifies the posture of the second user, and the third device identifies the posture of the third user.

[0075] Specifically, the second device can capture an image of the second user through a camera and recognize the second user's posture based on the image. The third device can capture an image of the third user through a camera and recognize the third user's posture based on the image. The posture of the second user can be used to determine the probability that the second user is speaking, and the posture of the third user can be used to determine the probability that the third user is speaking.

[0076] For example, the posture of the second user may include the following: the second user is not in the frame, the second user is in the frame, the second user is facing the camera, the second user is speaking to the camera, and the second user is waving to the camera. Similarly, the posture of the third user may also include the following: the third user is not in the frame, the third user is in the frame, the third user is facing the camera, the third user is speaking to the camera, and the third user is waving to the camera.

[0077] It is understood that the above user gestures are merely illustrative and do not constitute a limitation on the embodiments of this application. In some embodiments, other types of user gestures may also be included.

[0078] The posture recognition of the second and third users can be achieved through a preset image recognition algorithm. This application embodiment does not impose any special limitations on the preset image recognition algorithm.

[0079] In some alternative embodiments, the second device can continuously identify the posture of the second user, and the third device can continuously identify the posture of the third user.

[0080] In some optional embodiments, since posture recognition requires power consumption and computing power, in order to save power consumption and computing power for the second and third devices, the second device can periodically recognize the posture of the second user, and the third device can periodically recognize the posture of the third user. The period can be preset, or it can be manually adjusted by the user; this application embodiment does not impose any special limitations on this.

[0081] In some optional embodiments, in order to save power consumption and computing power of the second and third devices, the second device can identify the posture of the second user when it detects the second user's voice; the third device can identify the posture of the third user when it detects the third user's voice.

[0082] It is understood that the first device can also recognize the posture of the first user. The specific way in which the first device recognizes the posture of the first user can be referred to the way in which the second device recognizes the posture of the second user or the way in which the third device recognizes the posture of the third user in the above embodiments, which will not be repeated here.

[0083] In step 302, the second device sends the second user's posture to the first device, and the third device sends the third user's posture to the first device. Accordingly, the first device receives the postures of the second user and the third user.

[0084] Specifically, after recognizing the posture of the second user, the second device can send the posture of the second user to the first device; after recognizing the posture of the third user, the third device can send the posture of the third user to the first device.

[0085] It is understandable that, since the second device, the third device, and the first device are in a multi-party call scenario, the second device can send the second user's gestures to the first device and the third device at the same time, thereby enabling the third device to sort and display the video feeds of the participants; or, the third device can send the third user's gestures to the first device and the second device at the same time, thereby enabling the second device to sort and display the video feeds of the participants.

[0086] Similarly, the first device can also send the first user's posture to the second and third devices.

[0087] Understandably, when the second device sends the second user's gesture, it can also carry an identification number corresponding to the gesture. This identification number uniquely identifies the second device, allowing the first device, upon receiving the second user's gesture, to determine which device sent the gesture. Similarly, when the third device sends the third user's gesture, it can also carry an identification number corresponding to the gesture. This identification number uniquely identifies the third device, allowing the first device, upon receiving the third user's gesture, to determine which device sent the gesture.

[0088] Step 303: The first device determines the display positions of the second user's video screen and the third user's video screen on the multi-party video call interface of the first device based on the posture of the second user and the posture of the third user.

[0089] Specifically, after receiving the gestures of the second user sent by the second device and the gestures of the third user sent by the third device, the first device can determine the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the gestures of the second user and the third user. As one implementation, the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device can be determined according to the priority of the second user's video feed and the third user's video feed.

[0090] For example, assuming that the video feed of the second user has a higher priority than the video feed of the third user, the video feed of the second user with higher priority can be displayed on the homepage (i.e., the main interface) of the first device, while the video feed of the third user with lower priority can be displayed on other pages. As one implementation, these other pages are the pages displayed after the user swipes or performs other operations on the homepage. For example, the other pages can be the second page, the third page, etc.

[0091] It is understood that the above three pages are merely illustrative and do not constitute a limitation on the embodiments of this application. In some embodiments, more or fewer pages may be included.

[0092] The priority can be determined based on the posture of the second user and the posture of the third user.

[0093] For example, the postures of the second user include the second user not appearing in the frame, the second user appearing in the frame, the second user facing the camera, the second user speaking to the camera, and the second user waving to the camera; and the postures of the third user include the third user not appearing in the frame, the third user appearing in the frame, the third user facing the camera, the third user speaking to the camera, and the third user waving to the camera.

[0094] If the first device determines that the second user's gesture is waving at the camera, and / or the third user's gesture is that the third user is not visible in the frame, it can determine that the second user's video feed has a higher priority than the third user's video feed. In other words, the second user's video feed can be displayed first on the homepage, and the third user's video feed can be displayed on later pages, such as the second page, the third page, etc. Alternatively,

[0095] If the first device determines that the second user's posture is that the second user is not in the frame, and / or the third user's posture is that the third user is waving at the camera, it can determine that the third user's video feed has a higher priority than the second user's video feed. In other words, the third user's video feed can be displayed on the homepage first, and the second user's video feed can be displayed on a later page, such as the second page, the third page, etc.

[0096] It is understandable that after receiving the gestures of the users on the first and third devices, the second device can also determine the display positions of the first user's video feed and the third user's video feed on the multi-party video call interface of the second device. Alternatively, after receiving the gestures of the users on the first and second devices, the third device can also determine the display positions of the first user's video feed and the second user's video feed on the multi-party video call interface of the third device.

[0097] The method for determining the display positions of the first user's video feed and the third user's video feed on the multi-party video call interface of the second device, as well as the method for determining the display positions of the first user's video feed and the second user's video feed on the multi-party video call interface of the third device, can be found in the method for determining the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device, and will not be elaborated here.

[0098] In some optional embodiments, multiple users' screens can be displayed on a single page. These multiple screens can be screens of different sizes. For example, assuming the second user's video screen has a higher priority than the third user's video screen, the higher-priority second user's video screen can be displayed on a larger screen, and the lower-priority third user's video screen can be displayed on a smaller screen. Alternatively, assuming the third user's video screen has a higher priority than the second user's video screen, the higher-priority third user's video screen can be displayed on a larger screen, and the lower-priority second user's video screen can be displayed on a smaller screen.

[0099] Step 304: The first device displays the video feeds of the second user and the third user according to the determined display positions.

[0100] Figure 4a and Figure 4b A diagram illustrating the sorting and display of video frames.

[0101] Figure 4a This is a schematic diagram of one embodiment of a video frame sorting and display method. (Reference) Figure 4aAssuming the second user's video feed has a higher priority than the third user's video feed, the first device's homepage can display the higher-priority second user's video feed; for example, the homepage displays the video feed of second user A. The first device's second page can display the lower-priority third user's video feed; for example, the second page displays the video feed of third user B. The second page is the page displayed after the user swipes or performs other actions on the homepage.

[0102] Figure 4b A schematic diagram of another embodiment of a video frame sorting display method. (See also...) Figure 4b Assuming the second user's video feed has higher priority than the third user's, the higher-priority second user's video feed can be displayed in the first area of ​​the first device's homepage, while the lower-priority third user's video feed can be displayed in the second area of ​​the first device's homepage. The first area is larger than the second area. This allows the first user to view the user they are currently conversing with on the main screen, improving the user experience.

[0103] like Figure 5 The diagram shown is a flowchart of another embodiment of the call method provided in this application. Figure 5 In the illustrated embodiment, the first device can sort and display the video feeds of the participants based on the user's volume and / or speech content collected by the second device and the user's volume and / or speech content collected by the third device. Specifically, this includes the following steps:

[0104] Step 501: The second device collects the voice of the second user, and the third device collects the voice of the third user.

[0105] Specifically, when the second user is speaking, the second device can capture the second user's voice through a microphone; when the third user is speaking, the third device can capture the third user's voice through a microphone.

[0106] In step 502, the second device sends the second user's audio stream to the first device, and the third device sends the third user's audio stream to the first device. Correspondingly, the first device receives the second user's audio stream sent by the second device and the third user's audio stream sent by the third device.

[0107] Specifically, after the second device collects the voice of the second user, it can send the audio stream of the second user to the first device; after the third device collects the voice of the third user, it can send the audio stream of the third user to the first device.

[0108] Step 503: The first device acquires the volume and / or speech content of the second user based on the received audio stream of the second user, and acquires the volume and / or speech content of the third user based on the received audio stream of the third user.

[0109] Specifically, after receiving the audio stream of the second user sent by the second device, the first device can obtain the corresponding volume and / or speech content based on the received audio stream of the second user, and after receiving the audio stream of the third user sent by the third device, it can obtain the corresponding volume and / or speech content based on the received audio stream of the third user.

[0110] Volume can be used to represent the decibel level at which a user speaks.

[0111] For example, the first device may sample the audio stream of the second user once arbitrarily and use the sampled value as the volume of the second user; or, the first device may sample the audio stream of the third user arbitrarily once and use the sampled value as the volume of the third user.

[0112] In some optional embodiments, to reduce errors, the audio stream can be sampled multiple times, and the average of the multiple samples can be used as the volume of the corresponding user. For example, the first device can sample the audio stream of the second user multiple times and use the average of the multiple sample values ​​as the volume of the second user; or, the first device can sample the audio stream of the third user multiple times and use the average of the multiple sample values ​​as the volume of the third user.

[0113] It is understood that the spoken content can be obtained by performing speech recognition based on a preset speech recognition algorithm. This application does not specifically limit the preset speech recognition algorithm. After receiving the audio stream sent by the second or third device, the first device can recognize the spoken content of the second or third user based on the preset speech recognition algorithm.

[0114] Step 504: The first device determines the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the second user's volume and / or speech content, and the third user's volume and / or speech content.

[0115] Specifically, after the first device obtains the volume and / or speech content of the second user and the volume and / or speech content of the third user, it can determine the display positions of the video screens of the second user and the third user on the multi-party video call interface of the first device in the following manner.

[0116] For example, the first device can determine the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the volume of the second user and the volume of the third user.

[0117] After acquiring the volume data of the second and third users, the first device can prioritize the video feeds of the second and third users based on their volume. It can then determine the display position of their video feeds on the multi-party video call interface within the first device. For example, higher volume means higher priority; that is, the video feed of the user with the higher volume can be displayed on the homepage, while the video feed of the user with the lower volume can be displayed on other pages, such as the second or third page.

[0118] In some optional embodiments, multiple users' screens can be displayed on a single page. The specific method for displaying video feeds of users with different priorities on a single page can be found in the descriptions above, and will not be repeated here.

[0119] In some optional embodiments, the first device may determine the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the content of the second user's speech and the content of the third user's speech.

[0120] After acquiring the speech content of the second user and the third user, the first device can prioritize the video feeds of the second user and the third user based on their speech content, and determine the display positions of the video feeds of the second user and the third user on the multi-party video call interface of the first device based on their priorities.

[0121] One method for determining the priority of a user's video feed based on the content of the second and third users' speech is that the first device can determine whether it is calling the first user based on the content of the second and third users' speech. In this case, the user calling the first user has a higher priority, and the user not calling the first user has a lower priority.

[0122] For example, if a second user wants to speak, they can say, "Xiao Wu, let me share my thoughts." Xiao Wu can be the first user. In this case, the second user speaking has higher priority, and their video feed can be displayed on the homepage. Alternatively, the second user's video feed can be displayed on a larger screen on the homepage.

[0123] In some alternative embodiments, the first device may determine the display positions of the video feeds of the second user and the third user on the multi-party video call interface of the first device based on the content of the first user's speech.

[0124] After acquiring the first user's speech content, the first device can prioritize the video feeds of the second user and the third user based on the first user's speech content, and determine the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the priority of the second user's video feed and the third user's video feed.

[0125] The specific method by which the first device recognizes the first user's speech can be referenced from the method by which the first device recognizes the speech of the second or third user, and will not be repeated here. The method for determining the priority of user video feeds based on the first user's speech can be that the first device can determine the called party based on the first user's speech; users corresponding to the called party have higher priority, while users not corresponding to the called party have lower priority.

[0126] For example, when the first user actively calls out to the second user, Xiao Wang, they can say, "Xiao Wang, please speak." In this case, Xiao Wang, as the second user, has higher priority, and Xiao Wang's video can be displayed on the homepage. Alternatively, Xiao Wang's video can be displayed on a larger screen on the homepage.

[0127] In some optional embodiments, if the first device cannot determine a specific called party from the first user's speech, it can match the first user's speech with the speech of the second user and the third user to determine whether the second user's speech or the third user's speech has a higher match degree, so that the video of the user with the higher match degree is prioritized for display on the homepage. Alternatively, the video of the user with the higher match degree can be prioritized for display on a larger screen on the homepage.

[0128] The method for matching the speech content of the first user with the speech content of the second user and the speech content of the third user can be as follows: the semantics of the first user are identified by a preset semantic recognition algorithm. Then, the number of words with the same or similar semantics as the first user can be counted in the speech content of the second user and the speech content of the third user. The more words with the same or similar semantics as the first user, the more the speech content of the first user matches the speech content of the third user, which means that the user has a higher priority.

[0129] Alternatively, a preset semantic recognition algorithm can be used to identify the semantics of the first user, the second user, and the third user respectively. Then, the similarity between the semantics of the first user, the second user, and the third user can be calculated. The higher the similarity, the more the first user's speech content matches the second user's speech content, which means that the user has a higher priority.

[0130] It is understood that the above matching methods are merely illustrative and do not constitute a limitation on the embodiments of this application. In some embodiments, other matching methods may also be used.

[0131] In some optional embodiments, the first device may first determine the priority of the second user and the third user based on the volume of the second user and the third user. If the second user and the third user have the same priority, it may then determine which user's speech content matches the first user's speech content better. The user whose speech content matches better may have a higher priority than the user whose speech content does not match. The methods for determining priority based on volume and matching based on speech content can be found in the relevant descriptions in the above embodiments and will not be repeated here.

[0132] In some optional embodiments, the first device may first determine the priority of the second user and the third user based on the speech content of the first user, the speech content of the second user, and the speech content of the third user. If the second user and the third user have the same priority, it may determine which user has a louder volume. The user with a louder volume may have a higher priority than the user with a lower volume.

[0133] Step 505: The first device displays the video feeds of the second user and the third user according to the determined display positions.

[0134] Specifically, the display method of the video image can be referred to the relevant description in the above embodiments, and will not be repeated here.

[0135] Figure 6 This is a schematic diagram of the structure of one embodiment of the calling device of this application, as shown below. Figure 6 As shown, the aforementioned communication device 60 is applied to a first device, which conducts multi-party video calls with a second and a third device. The communication device 60 may include: an acquisition module 61, a determination module 62, and a display module 63; wherein,

[0136] Acquisition module 61 is used to acquire first information of the second user corresponding to the second device, the first information including at least one of audio information or video information; and acquire second information of the third user corresponding to the third device, the second information including at least one of audio information or video information;

[0137] The determining module 62 is used to determine the display positions of the video frames of the second user and the third user on the multi-party video call interface in the first device based on the first information and the second information.

[0138] Display module 63 is used to display the video screen of the second user and the video screen of the third user according to the determined display position.

[0139] In one possible implementation, the determining module 62 is further configured to determine the priority of the second user's video frame and the priority of the third user's video frame based on the first information and the second information;

[0140] The display positions of the video feeds of the second user and the third user on the multi-party video call interface of the first device are determined based on the priority of the video feeds of the second user and the third user.

[0141] In one possible implementation, the determining module 62 is further configured to, if the priority of the second user's video feed is higher than the priority of the third user's video feed, display the second user's video feed on the home page of the multi-party video call interface on the first device, and display the third user's video feed on other pages of the multi-party video call interface; or...

[0142] If the video feed of the second user has a higher priority than the video feed of the third user, the video feed of the second user will be displayed in the first area of ​​the homepage of the multi-party video call interface on the first device, and the video feed of the third user will be displayed in the second area. The homepage includes the first area and the second area, and the size of the first area is larger than the size of the second area.

[0143] In one possible implementation, the video information in the first information includes the posture of the second user, and the video information in the second information includes the posture of the third user; or...

[0144] The audio information in the first information includes the volume of the second user, and the audio information in the second information includes the volume of the third user.

[0145] In one possible implementation, the volume of the second user is determined by multiple sample values ​​in the audio stream of the second user acquired by the second device, and the volume of the third user is determined by multiple sample values ​​in the audio stream of the third user acquired by the third device.

[0146] In one possible implementation, the audio information in the first information includes the speech content of the second user, and the audio information in the second information includes the speech content of the third user.

[0147] In one possible implementation, the determining module 62 is further configured to determine whether the speech content of the second user and the speech content of the third user contain the first user corresponding to the first device;

[0148] If the second user's speech contains information about the first user, but the third user's speech does not contain information about the first user, the second user's video feed has a higher priority than the third user's video feed; or...

[0149] If the third user's speech includes the first user, but the second user's speech does not include the first user, the third user's video feed has a higher priority than the second user's video feed.

[0150] In one possible implementation, the determining module 62 is further configured to match the speech content of the first user corresponding to the first device with the speech content of the second user and the speech content of the third user, respectively;

[0151] If the match between the first user's speech and the second user's speech is higher than the match between the first user's speech and the third user's speech, then the second user's video feed has a higher priority than the third user's video feed; or...

[0152] If the matching degree between the first user's speech and the third user's speech is higher than the matching degree between the first user's speech and the second user's speech, the priority of the third user's video feed is higher than the priority of the second user's video feed.

[0153] In one possible implementation, the determining module 62 is further configured to:

[0154] Determine the target object in the speech content of the first user corresponding to the first device;

[0155] If the target of the first user's speech is the second user, the second user's video feed has a higher priority than the third user's video feed; or,

[0156] If the target of the first user's speech is the third user, the video feed of the third user has a higher priority than the video feed of the second user.

[0157] Figure 6 The communication device 60 provided in the illustrated embodiment can be used to execute the technical solution of the method embodiment shown in this application, and its implementation principle and technical effect can be further referred to the relevant description in the method embodiment.

[0158] The above should be understood Figure 6 The division of the various modules in the illustrated communication device 60 is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. For example, the detection module can be a separate processing element or integrated into a chip within the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0159] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, these modules can be integrated together as a System-On-a-Chip (SOC).

[0160] In the above embodiments, the processor may include, for example, a CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, embedded neural network processing unit (NPU), and image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as an ASIC, or one or more integrated circuits for controlling the execution of the program in this application. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium.

[0161] This application also provides a readable storage medium storing a program that, when run on a device, causes the device to execute the method provided in the embodiments shown in this application.

[0162] This application also provides a program that, when run on a device, causes the device to perform the method provided in the embodiments shown in this application.

[0163] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0164] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0165] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0166] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for making a call, characterized in that, The method, applied to a first device, wherein the first device engages in multi-party video calls with a second and a third device, includes: Obtain first information of the second user corresponding to the second device, the first information including audio information, the audio information in the first information including the speech content of the second user; Obtain second information of the third user corresponding to the third device, the second information including audio information, the audio information in the second information including the speech content of the third user; Based on the first information and the second information, determine the display positions of the video screens of the second user and the third user on the multi-party video call interface in the first device; The video feeds of the second user and the third user are displayed according to the determined display positions; Determining the display positions of the second user's video feed and the third user's video feed on the multi-party video call interface of the first device based on the first information and the second information includes: The priority of the second user's video feed and the priority of the third user's video feed are determined based on the first information and the second information. The display positions of the video frames of the second user and the third user on the multi-party video call interface in the first device are determined based on the priority of the video frames of the second user and the third user. Determining the priority of the second user's video feed and the priority of the third user's video feed based on the first information and the second information includes: If the recipient cannot be determined based on the speech content of the first user corresponding to the first device, then... The speech content of the first user corresponding to the first device is matched with the speech content of the second user and the speech content of the third user, respectively; If the matching degree between the first user's speech and the second user's speech is higher than the matching degree between the first user's speech and the third user's speech, then the priority of the second user's video feed is higher than the priority of the third user's video feed; or, If the matching degree between the first user's speech and the third user's speech is higher than the matching degree between the first user's speech and the second user's speech, the video frame of the third user has a higher priority than the video frame of the second user.

2. The method according to claim 1, characterized in that, The step of determining the display positions of the video frames of the second user and the third user on the multi-party video call interface in the first device based on the priority of the video frames of the second user and the third user includes: If the video feed of the second user has a higher priority than the video feed of the third user, the video feed of the second user will be displayed on the home page of the multi-party video call interface on the first device, and the video feed of the third user will be displayed on other pages of the multi-party video call interface; or... If the video feed of the second user has a higher priority than the video feed of the third user, the video feed of the second user will be displayed in the first area of ​​the homepage of the multi-party video call interface on the first device, and the video feed of the third user will be displayed in the second area. The homepage includes the first area and the second area, and the size of the first area is larger than the size of the second area.

3. The method according to claim 1 or 2, characterized in that, The first information also includes video information, and the second information also includes video information; The video information in the first information includes the posture of the second user, and the video information in the second information includes the posture of the third user; or, The audio information in the first information also includes the volume of the second user, and the audio information in the second information also includes the volume of the third user.

4. The method according to claim 3, characterized in that, The volume of the second user is determined by multiple sample values ​​in the audio stream of the second user acquired by the second device, and the volume of the third user is determined by multiple sample values ​​in the audio stream of the third user acquired by the third device.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Determine the target object in the speech content of the first user corresponding to the first device; If the target of the first user's speech is the second user, the second user's video feed has a higher priority than the third user's video feed; or, If the target of the first user's speech is the third user, the video feed of the third user has a higher priority than the video feed of the second user.

6. A device, characterized in that, include: Processor and memory, the memory being used to store programs; The processor is used to run the program to implement the call method as described in any one of claims 1-5.

7. A readable storage medium, characterized in that, The readable storage medium stores a program that, when run on the device, implements the call method as described in any one of claims 1-5.