An interactive method, device and system for audio and video call scenarios

By judging and switching the background environment of both parties during audio and video calls, the problem of poor user experience is solved, and immersive interaction and emotional communication are improved.

CN114979544BActive Publication Date: 2025-10-03SHENZHEN KONKA ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210641932.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-10-03
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

Existing audio and video call technologies do not consider the subjective impact of the environment in which the two parties in the call are located on the user experience, resulting in a poor user experience.

Method used

During an audio or video call, determine whether to switch the background environment of both parties. If necessary, switch them to the target scene. Otherwise, keep their respective background environments and use the audio or video call terminal, application server, and scene server to switch and display the scenes.

Benefits of technology

It improves the user experience of audio and video calls, realizes immersive interaction, and enhances emotional communication and interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114979544B_ABST
    Figure CN114979544B_ABST
Patent Text Reader

Abstract

The present disclosure provides an interactive method, device, and system for audio and video call scenarios, wherein the method includes: during an audio and video call, determining whether to switch the call scenario, wherein the call scenario is the background environment of both parties; if it is determined that the call scenario is to be switched, switching the background environment of both parties to the call to the target scenario; and if it is determined that the call scenario is not to be switched, maintaining the background environment of each party. Through this disclosure, the problem of poor user experience caused by the failure to consider the subjective influence of the environment on the two parties in the call in the related art of audio and video calls is solved, thereby achieving the purpose of improving the user experience of audio and video calls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an interactive method, device, and system for audio and video call scenarios. Background Art

[0002] With the continuous development of society, many people have left their hometowns for work, potentially away from their families, friends, and relatives for extended periods of time. Most communication methods rely on online audio and video calls. Current audio and video calls tend to focus on ensuring call quality or adding features that enhance the appearance of the caller, but rarely consider the subjective impact of the surroundings on both sides of the call, resulting in a poor user experience and a lack of immersive interaction.

[0003] Currently, no effective solution has been proposed to the problem of poor user experience caused by audio and video calls in related technologies because the subjective influence of the environment on the two parties in the call is not taken into consideration. Summary of the Invention

[0004] The purpose of the present disclosure is to address the deficiencies in the prior art and to provide an interactive method, device, and system for audio and video call scenarios, as well as an electronic device and a computer-readable storage medium, so as to at least solve the problem in the related art that audio and video calls have a poor user experience due to the failure to consider the subjective impact of the environment on the two parties in the call.

[0005] According to one aspect of the present disclosure, a method for interacting in an audio and video call scenario is provided, comprising:

[0006] During an audio or video call, determining whether to switch the call scene, wherein the call scene refers to the background environment of the two parties in the call;

[0007] If it is determined that the call scene needs to be switched, the background environment of both parties of the call is switched to the target scene;

[0008] If it is determined that the call scene does not need to be switched, the background environments of the two parties in the call are maintained.

[0009] According to another aspect of the present disclosure, an interactive device for an audio and video call scenario is provided, comprising:

[0010] A determination unit, configured to determine whether to switch a call scene during an audio or video call, wherein the call scene refers to the background environment of both parties of the call;

[0011] A switching unit, configured to switch the background environments of both parties of the call to the target scene if it is determined that the call scene needs to be switched;

[0012] The holding unit is used to maintain the background environment of each of the two parties in the call if it is determined that the call scene is not to be switched.

[0013] According to another aspect of the present disclosure, an interactive system for an audio and video call scenario is provided, comprising:

[0014] An audio and video call terminal device, comprising an audio and video call module and a call scene imaging module, wherein the call scene imaging module is used to perform imaging display based on call scene data;

[0015] an application server, communicatively connected to the audio and video call terminal device, comprising an audio and video data service unit and an operation instruction processing unit, wherein the operation instruction processing unit is configured to send a scene switching instruction to the scene server and receive the call scene data returned by the scene server;

[0016] The scene server is in communication with the application server and is configured to generate the call scene data according to the scene switching instruction and return the call scene data to the application server.

[0017] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0018] processor; and

[0019] Memory for storing programs,

[0020] The program includes instructions, which, when executed by the processor, enable the processor to execute the interactive method of the video call scenario in the present disclosure.

[0021] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the interactive method of the video call scenario in the present disclosure.

[0022] One or more technical solutions provided in the embodiments of the present disclosure determine whether to switch the call scene during an audio or video call, wherein the call scene is the background environment of both parties on the call; if it is determined that the call scene is to be switched, the background environment of both parties on the call is switched to the target scene; if it is determined that the call scene is not to be switched, the background environment of each party on the call is maintained. This disclosure solves the problem in the related art that audio and video calls have a poor user experience due to not considering the subjective influence of the environment on the two parties on the call, thereby achieving the effect of improving the user experience of audio and video calls. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0024] Figure 1 A schematic diagram illustrating an interactive system for an audio and video call scenario in which various methods described herein may be implemented according to an exemplary embodiment of the present disclosure is shown;

[0025] Figure 2 A flowchart illustrating an interactive method for an audio and video call scenario according to an exemplary embodiment of the present disclosure is shown;

[0026] Figure 3 A flowchart illustrating an optional interactive method for an audio and video call scenario according to an exemplary embodiment of the present disclosure is shown;

[0027] Figure 4 A schematic block diagram of an interactive device for an audio and video call scenario according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0029] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0030] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0031] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] Aspects of the present disclosure are described below with reference to the accompanying drawings.

[0035] An exemplary embodiment of the present disclosure provides an interactive system for audio and video call scenarios. Figure 1 A schematic diagram of an interactive system for an audio and video call scenario in which various methods described herein can be implemented according to an exemplary embodiment of the present disclosure is shown. Figure 1 As shown, the system includes: an audio and video call terminal device 10, an application server 20 and a scene server 30.

[0036] The audio and video call terminal device 10 includes an audio and video call module 101 and a call scene imaging module 102, wherein the audio and video call module 101 is used to control the audio and video calls between the two parties of the call, and the call scene imaging module 102 is used to perform imaging display according to the call scene data.

[0037] In some embodiments, the call scene imaging module 202 may include:

[0038] The image processing module is used to collect the background environment of both parties in the call, such as cameras and other devices that can capture images.

[0039] The imaging projection module is used to perform imaging display based on the call scene data. The imaging projection module can display the final call scene in 3D to form an immersive call environment.

[0040] The application server 20 is communicatively connected to the audio and video call terminal device 10, and includes an audio and video data service unit 201 and an operation instruction processing unit 202, wherein the audio and video data service unit 201 is used to control the audio and video calls of the two parties and process the audio and video data of the two parties, and the operation instruction processing unit 202 is used to send the scene switching instruction to the scene server 30 and receive the call scene data returned by the scene server 30.

[0041] The scenario server 30 is in communication with the application server 20 and includes a machine learning modeling module 301 and a scenario generation service module 302. The machine learning modeling module 301 is configured to generate a call scenario model based on the scenario switching instruction, and the scenario generation service module 302 is configured to generate the call scenario data based on the call scenario model. After generating the call scenario data based on the scenario switching instruction, the scenario server 30 returns the call scenario data to the application server 20.

[0042] It should be noted that the audio and video call scene in the exemplary embodiments of the present disclosure refers to the physical environment in which the two parties to the call are located, that is, the background environment during the call.

[0043] In an exemplary embodiment of the present disclosure, an audio and video call terminal device communicates with an application server in the cloud through user operation instructions; the application server conducts a normal audio and video call according to the corresponding instructions, and completes the generation of the current call scene through communication with the scene server through scene switching instructions; after receiving the specific scene switching instruction, the scene server distinguishes the scene type, obtains basic scene data, and generates a corresponding call scene model. After the call scene model is created, the scene generation service module generates specific call scene data, and then notifies the operation instruction processing unit of the application server, and the operation instruction processing unit then notifies the audio and video call terminal device to complete the call scene imaging display.

[0044] An exemplary embodiment of the present disclosure provides an interactive method for an audio and video call scenario. Figure 2 A flowchart of an interactive method for an audio and video call scenario according to an exemplary embodiment of the present disclosure is shown. Figure 2 As shown, the method includes the following steps:

[0045] Step S201, during an audio or video call, determining whether to switch a call scene, wherein the call scene refers to the background environment of both parties of the call;

[0046] Step S202: If it is determined that the call scene is to be switched, the background environment of both parties of the call is switched to the target scene;

[0047] Step S203: If it is determined that the call scene is not to be switched, the background environments of the two call parties are maintained.

[0048] Through the above steps, the problem of poor user experience caused by audio and video calls in related technologies due to failure to consider the subjective influence of the environment on the two parties in the call is solved, and the effect of improving the user experience of audio and video calls is achieved.

[0049] In some embodiments, during an audio or video call, a call scene switching operation touch area may be set in the call interface. If a touch signal in the call scene switching operation touch area is detected, it may be determined that the call scene is switched.

[0050] If it is determined that the call scene is not to be switched, the background environments of the two call parties can be maintained; if it is determined that the call scene is to be switched, the background environments of both call parties are switched to the target scene.

[0051] In some embodiments, if it is determined that the call scene needs to be switched, the background environment of both parties of the call is switched to the target scene, including:

[0052] generating a call scenario model of the target scenario;

[0053] Generating call scenario data of the call scenario model;

[0054] The call scene data is sent to the audio and video call terminal devices of both parties of the call for display.

[0055] In some embodiments, both parties of a call can choose to switch the call scene to the system scene, or can choose to switch the call scene to the background environment of either party.

[0056] If it is determined that the call scene is switched to one of the system scenes, one of the system scenes can be determined as the target scene. In this case, a call scene model of the target scene can be directly generated based on one of the system scenes; call scene data of the call scene model can then be generated; and finally, the call scene data can be sent to the audio and video call terminal devices of both parties for display.

[0057] When it is determined that the call scene is switched to the background environment of one of the two parties in the call, the background environment data of one of the two parties in the call is first collected, and a call scene model of the target scene is generated based on the background environment data; then the call scene data of the call scene model is generated; finally, the call scene data is sent to the audio and video call terminal devices of both parties in the call for display.

[0058] The exemplary embodiments of the present disclosure provide an interactive method for an immersive audio and video call scenario. The specific interactive process can be as follows: Figure 3 As shown, the process includes the following steps S301 to S311.

[0059] In step S301, the user opens an application with audio and video calling function on a device such as a smart TV or a mobile phone, and makes an audio and video call according to a dialing rule.

[0060] In step S302, the called party connects the audio and video call and can see the caller's screen and hear the audio normally. This is a normal call scenario, which follows the current environment scenario of each party.

[0061] In step S303, the two parties negotiate the desired call scenario and select an appropriate immersive call scenario using the function buttons provided in the application interface. Call scenario types may include: scenarios based on the caller's current location, such as a gathering of friends; scenarios based on the called party's current location, such as a parent wanting to see their child's living environment; and system-supported scenarios, such as famous tourist attractions.

[0062] Step S304: Determine whether to switch the call scene. If step S305 is executed and the current call scene is not switched, the current normal call scene is continued. If a certain scene is selected, step S306 is executed to switch the current call scene according to the scene rule.

[0063] In step S307, if the current environment of the called party or the calling party is selected as a call scene, step S308 is executed to obtain the selected environmental scene data through the camera, and generate a call scene model through the cloud server through the artificial intelligence deep learning algorithm.

[0064] In step S309, if the system scenario is selected, a list of scenarios will pop up on the application interface for the user to select. If a famous scenic spot mentioned in the current call is available in the system scenario, the user can choose to switch. Then, step S310 is executed to generate a call scenario model based on the deep learning algorithm on the cloud server according to the set scenario;

[0065] Step S311: Through 3D imaging, virtual reality, and artificial intelligence technologies, the call scene is projected in front of both parties of the call through the imaging projection module, forming an immersive face-to-face call experience in the same environment.

[0066] The present disclosure provides users with an immersive audio and video call interaction method based on APP, cloud platform, and smart terminal devices, based on audio and video calls, artificial intelligence, virtual reality, 3D imaging technology, and immersive multi-scene switching to simulate a real environment. The new call scene interaction method highlights the immersive and intelligent use, enables users to get greater interaction, enhances the emotional communication between the two parties of the call, further improves the experience of audio and video calls, optimizes the interaction process, and deeply integrates multiple technologies based on user usage scenarios, reflecting the high-tech sense and humanistic care of the product, highlighting that technology is people-oriented, and enhancing emotional communication and user experience.

[0067] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0068] The exemplary embodiments of the present disclosure also provide an interactive device for audio and video call scenarios, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated. As used below, the terms "module," "unit," "subunit," etc. may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0069] Figure 4 A schematic block diagram of an interactive device for an audio and video call scenario according to an exemplary embodiment of the present disclosure is shown. Figure 4 As shown, the device includes:

[0070] The judgment unit 41 is used to judge whether to switch the call scene during the audio and video call, wherein the call scene is the background environment of the two parties in the call;

[0071] The switching unit 42 is configured to switch the background environment of both parties of the call to the target scene if it is determined that the call scene needs to be switched;

[0072] The maintaining unit 43 is configured to maintain the background environments of both parties in the call if it is determined that the call scene is not to be switched.

[0073] In some embodiments, the switching unit 42 includes:

[0074] A first generating module, configured to generate a call scenario model of the target scenario;

[0075] A second generating module, configured to generate call scenario data of the call scenario model;

[0076] The display module is used to send the call scene data to the audio and video call terminal devices of both parties of the call for display.

[0077] In some embodiments, the switching unit 42 includes:

[0078] A determination module is configured to, when it is determined that the call scenario is to be switched to one of the system scenarios, determine one of the system scenarios as the target scenario.

[0079] In some embodiments, the first generating module includes:

[0080] A collection submodule, configured to collect background environment data of one of the two call parties when it is determined that the call scene is switched to the background environment of one of the two call parties;

[0081] A generation submodule is used to generate a call scene model of the target scene according to the background environment data.

[0082] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0083] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.

[0084] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.

[0085] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.

[0086] refer to Figure 5 , a block diagram of an electronic device 500 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0087] like Figure 5As shown, electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of device 500 can also be stored in RAM 503. Computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.

[0088] Multiple components within electronic device 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into electronic device 500. Input unit 506 can receive input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0089] The computing unit 501 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the interactive method of the audio and video call scenario can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to execute the interactive method of the audio and video call scenario by any other appropriate means (for example, by means of firmware).

[0090] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0091] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0092] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0093] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0094] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0095] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. An interactive method for audio and video call scenarios, characterized in that: include: During an audio or video call, determining whether to switch the call scene, wherein the call scene refers to the background environment of the two parties in the call; If it is determined that the call scene needs to be switched, the background environment of both parties of the call is switched to the target scene; If it is determined that the call scene does not need to be switched, the background environment of the two parties in the call is maintained; If it is determined that the call scene is to be switched, the background environment of both parties of the call is switched to the target scene, including: generating a call scenario model of the target scenario; Generate call scenario data of the call scenario model, wherein the call scenario model and the call scenario data are generated by the scenario server and fed back to the application server; The call scene data is sent to the audio and video call terminal devices of the call parties for display, wherein the imaging projection module in the audio and video call terminal device projects the call scene in front of the call parties for 3D imaging display based on the call scene data using 3D imaging, virtual reality, and artificial intelligence technologies; the call scene data is fed back to the audio and video call terminal by the application server; The step of generating the call scenario model of the target scenario includes: In the case where it is determined that the call scene is switched to the background environment of one of the two call parties, collecting background environment data of one of the two call parties; A call scene model of the target scene is generated according to the background environment data.

2. The interactive method for audio and video call scenarios according to claim 1, characterized in that: If it is determined that the call scene needs to be switched, the background environment of both parties in the call will be switched to the target scene, including: In the case where it is determined to switch the call scene to one of the system scenes, one of the system scenes is determined as the target scene.

3. An interactive device for audio and video call scenarios, characterized in that: include: A determination unit, configured to determine whether to switch a call scene during an audio or video call, wherein the call scene refers to the background environment of both parties of the call; A switching unit, configured to switch the background environments of both parties of the call to the target scene if it is determined that the call scene needs to be switched; A holding unit, configured to maintain the background environments of both parties in the call if it is determined that the call scene is not to be switched; Wherein, the switching unit is further used for: generating a call scenario model of the target scenario; Generating call scenario data of the call scenario model; The call scene data is sent to the audio and video call terminal devices of the call parties for display, wherein the imaging projection module in the audio and video call terminal devices uses 3D imaging, virtual reality, and artificial intelligence technologies to project the call scene in front of the call parties for 3D imaging display based on the call scene data; Wherein, the switching unit is further used for: In the case where it is determined that the call scene is switched to the background environment of one of the two call parties, collecting background environment data of one of the two call parties; A call scene model of the target scene is generated according to the background environment data.

4. An interactive system for audio and video call scenarios, characterized in that: include: An audio and video call terminal device, comprising an audio and video call module and a call scene imaging module, wherein the call scene imaging module is used to project the call scene in front of the call parties for 3D imaging display based on the call scene data using 3D imaging, virtual reality, and artificial intelligence technologies; an application server, communicatively connected to the audio and video call terminal device, comprising an audio and video data service unit and an operation instruction processing unit, wherein the operation instruction processing unit is configured to send a scene switching instruction to the scene server and receive the call scene data returned by the scene server; The scene server is communicatively connected to the application server, and is configured to generate a call scene model and the call scene data according to the scene switching instruction, and return the call scene data to the application server; The call scene imaging module is further configured to collect background environment data of one of the two call parties when it is determined that the call scene is switched to the background environment of one of the two call parties; The scenario server is further configured to generate a call scenario model of a target scenario based on the background environment data.

5. The interactive system for a video call scenario according to claim 4, wherein: The call scene imaging module includes: Image processing module, used to collect the background environment of both parties in the call; The imaging projection module is used to perform imaging display according to the call scene data.

6. The interactive system for a video call scenario according to claim 5, wherein: The scene server includes: A machine learning modeling module, configured to generate a call scenario model according to the scenario switching instruction; The scenario generation service module is used to generate the call scenario data according to the call scenario model.

7. An electronic device, characterized in that: include: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to execute the interactive method for the video call scenario according to claim 1 or 2.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the interactive method for the video call scenario according to claim 1 or 2.

Citation Information

Patent Citations

  • Video call method and device, terminal and storage medium

    CN113411537A

  • Video communication method, apparatus and system

    WO2009152769A1

  • Interaction method and apparatus for video call

    WO2022089273A1