Video interaction method, apparatus and electronic device

By generating and displaying 3D models of target objects during video interaction, the limitations and singularity of existing video interaction methods are solved, thereby improving user experience and interactivity.

CN116614595BActive Publication Date: 2026-04-14SPREADTRUM COMM (TIANJIN) INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SPREADTRUM COMM (TIANJIN) INC
Filing Date
2023-06-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing video interaction methods are limited and monotonous, resulting in a poor user experience and failing to fully enhance the sense of interaction.

Method used

The system uses 3D reconstruction technology to display a 3D model of a target object in a video interactive scene. It generates a 3D model of the target object using a sequence of video frames and displays it in the video interactive screen. It also supports augmented reality function buttons to control the display and operation of the 3D model.

Benefits of technology

It improves the user experience during video interaction, enhances the sense of interaction, and provides a more intuitive and three-dimensional way of displaying target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116614595B_ABST
    Figure CN116614595B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of communication, and particularly relate to a video interaction method, device and electronic equipment. The video interaction method is applied to a first terminal side, and includes: obtaining a video frame sequence of a first target object, the video frame sequence including video frames of each shooting angle of the first target object; sending the video frame sequence to a second terminal side, the video frame sequence being used to generate a three-dimensional model of the first target object at the second terminal side; or, obtaining the video frame sequence of the first target object, generating a three-dimensional model of the first target object according to the video frame sequence; and sending the three-dimensional model of the first target object to the second terminal side. According to the embodiments of the present application, the three-dimensional model of the target object can be displayed in a video interaction scene through three-dimensional reconstruction, so that the interaction between the two parties in the video interaction can be further enhanced, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a video interaction method, apparatus and electronic device. Background Technology

[0002] With the development of internet technology, video interaction has gradually emerged. In video interaction, the two parties primarily exchange information through text, audio, and video. While this method offers users intuitiveness and immediacy, it also suffers from certain limitations and a degree of singularity. Therefore, enhancing the interactive experience for both parties in a video interaction has become a problem that needs to be addressed. Summary of the Invention

[0003] This invention provides a video interaction method, device, and electronic device. Through 3D reconstruction, a 3D model of a target object can be displayed in a video interaction scene, thereby further enhancing the sense of interaction between the two parties in the video interaction and improving the user experience.

[0004] In a first aspect, embodiments of the present invention provide a video interaction method, the method being applied to a first terminal side, comprising:

[0005] Obtain a video frame sequence of a first target object, the video frame sequence containing video frames of the first target object from various shooting angles; send the video frame sequence to a second terminal, the video frame sequence being used to generate a 3D model of the first target object on the second terminal; or,

[0006] Obtain the video frame sequence of the first target object, generate a three-dimensional model of the first target object based on the video frame sequence, and send the three-dimensional model of the first target object to the second terminal side.

[0007] In one possible implementation, obtaining the video frame sequence of the first target object includes:

[0008] The first terminal displays a first video interaction screen, which includes a first augmented reality (AR) function button.

[0009] After detecting the first trigger action of the first AR function button, the first AR mode is activated, wherein the video frames captured in the first AR mode are a sequence of video frames about the first target object.

[0010] In one possible implementation, after activating the first AR mode, the method further includes:

[0011] Send a first instruction to the second terminal side, the first instruction being used to instruct the second terminal side to determine a video frame sequence about the first target object from the received video frames.

[0012] In one possible implementation, sending the video frame sequence to the second terminal includes:

[0013] In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the video frame sequence is sent to the second terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or,

[0014] If the IP Multimedia Subsystem data channel is not supported between the first terminal and the second terminal, the video frame sequence is obtained and then sent to the second terminal.

[0015] In one possible implementation, generating the 3D model of the first target object based on the video frame sequence includes:

[0016] In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, a three-dimensional model of the first target object is generated based on the video frame sequence, and the three-dimensional model of the first target object is sent to the second terminal through the IP Multimedia Subsystem data channel.

[0017] In one possible implementation, after generating the 3D model of the first target object based on the video frame sequence, the method further includes:

[0018] A three-dimensional model of the first target object is displayed in the first video interactive screen on the first terminal side.

[0019] In one possible implementation, the method further includes:

[0020] When a first operation command is detected that acts on the three-dimensional model of the first target object in the first video interaction screen, the display effect of the three-dimensional model of the first target object is adjusted according to the first operation command.

[0021] The first operation instruction is sent to the second terminal side, and the second terminal side displays a second video interaction screen. The second video interaction screen displays a three-dimensional model of the first target object. The first operation instruction is also used by the second terminal side to adjust the display effect of the three-dimensional model in the second video interaction screen.

[0022] In one possible implementation, the method further includes:

[0023] Receive a second operation command sent by the second terminal side, the second operation command acting on the three-dimensional model in the second video interaction screen of the second terminal side;

[0024] Adjust the display effect of the three-dimensional model of the first target object in the first video interaction screen according to the second operation instruction.

[0025] Secondly, embodiments of the present invention provide a video interaction method, the method being applied to a second terminal side, comprising:

[0026] Receive a sequence of video frames about a first target object sent by a first terminal, generate a three-dimensional model of the first target object based on the video frame sequence, and display the three-dimensional model of the first target object in a second video interactive screen on a second terminal; or,

[0027] The system receives a 3D model of a first target object sent from the first terminal side; and displays the 3D model of the first target object in a second video interactive screen on the second terminal side.

[0028] In one possible implementation, receiving the video frame sequence about the first target object sent by the first terminal includes:

[0029] The second video interaction screen on the second terminal side includes a second AR function button; after detecting a second trigger action on the second AR function button, a second AR mode is activated, wherein the video frames received from the first terminal side in the second AR mode are a sequence of video frames about the first target object; or...

[0030] After receiving the first instruction sent by the first terminal, a video frame sequence about the first target object is determined from the video frames sent by the first terminal according to the first instruction.

[0031] In one possible implementation, receiving the video frame sequence about the first target object sent by the first terminal includes:

[0032] In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the system receives a video frame sequence about the first target object sent by the first terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or...

[0033] In response to the lack of support for an IP Multimedia Subsystem data channel between the second terminal and the first terminal, the system receives a sequence of video frames about the first target object sent by the first terminal.

[0034] In one possible implementation, receiving the three-dimensional model of the first target object sent by the first terminal side includes:

[0035] In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, the device receives a three-dimensional model of the first target object sent by the first terminal through the IP Multimedia Subsystem data channel.

[0036] In one possible implementation, after receiving the three-dimensional model of the first target object sent by the first terminal side, the method further includes:

[0037] Receive the first operation instruction sent by the first terminal side;

[0038] Adjust the display effect of the three-dimensional model of the first target object in the second video interaction screen according to the first operation instruction.

[0039] In one possible implementation, after receiving the three-dimensional model of the first target object sent by the first terminal side, the method further includes:

[0040] When a second operation command is detected that acts on the 3D model of the first target object in the second video interaction screen, the second operation command is sent to the first terminal side. The second operation command is also used by the first terminal side to adjust the display effect of the 3D model of the first target object in the first video interaction screen.

[0041] In one possible implementation, the method further includes: when a second operation command is detected acting on the three-dimensional model of the first target object in the second video interaction screen, adjusting the display effect of the three-dimensional model of the first target object according to the second operation command.

[0042] Thirdly, embodiments of the present invention provide a video interaction device, comprising:

[0043] The first transmission module is used to send a video frame sequence of the first target object to the second terminal side, and the video frame sequence of the first target object is used to generate a three-dimensional model of the first target object on the second terminal side.

[0044] or,

[0045] The video interaction device further includes: a first generation module, configured to generate a three-dimensional model of the first target object based on a video frame sequence of the first target object; and a first transmission module, configured to send the three-dimensional model of the first target object to the second terminal side.

[0046] Fourthly, embodiments of the present invention provide a video interaction device, comprising:

[0047] The second acquisition module is used to receive a video frame sequence of the first target object sent by the first terminal side;

[0048] The second generation module is used to generate a three-dimensional model of the first target object based on the video frame sequence of the first target object;

[0049] The second display module is used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side;

[0050] or,

[0051] The second transmission module is used to receive the three-dimensional model of the first target object sent by the first terminal side; the second display module is used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side.

[0052] Fifthly, embodiments of the present invention provide an electronic device, comprising:

[0053] At least one processor; and

[0054] At least one memory communicatively connected to the processor, wherein:

[0055] The memory stores program instructions that can be executed by the processor, and the processor can execute the method embodiment provided in either the first aspect or the second aspect by calling the program instructions.

[0056] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium storing computer instructions that cause the computer to perform the method embodiments provided in either the first or second aspect.

[0057] The present invention supports the reconstruction and display of a 3D model based on video frame data of a target object during video interaction. Simultaneously, the 3D model can be transmitted along with the audio and video data to the other terminal for display, facilitating user communication. Attached Figure Description

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a schematic diagram of a video interaction scenario provided in an embodiment of the present invention;

[0060] Figure 2 A flowchart of a video interaction method provided in an embodiment of the present invention;

[0061] Figure 3 A flowchart of another video interaction method provided in an embodiment of the present invention;

[0062] Figure 4 A flowchart illustrating another video interaction method provided in an embodiment of the present invention;

[0063] Figure 5 This is a schematic diagram illustrating another video interaction scenario provided by an embodiment of the present invention;

[0064] Figure 6 A flowchart illustrating yet another video interaction method provided in an embodiment of the present invention;

[0065] Figure 7 This diagram illustrates yet another video interaction scenario for an embodiment of the present invention.

[0066] Figure 8 A flowchart of an operation instruction interaction method provided in an embodiment of the present invention;

[0067] Figure 9 This is a schematic diagram illustrating the adjustment of the display effect of a three-dimensional model according to an embodiment of the present invention;

[0068] Figure 10 A flowchart of a three-dimensional model construction method provided in an embodiment of the present invention;

[0069] Figure 11 This is a schematic diagram of the structure of a video interaction device provided in an embodiment of the present invention;

[0070] Figure 12This is a schematic diagram of another video interaction device provided in an embodiment of the present invention;

[0071] Figure 13 This is a schematic diagram of the structure of a video interaction device provided in an embodiment of the present invention;

[0072] Figure 14 A schematic diagram of another video interaction device provided in an embodiment of the present invention;

[0073] Figure 15 This is a schematic diagram of the structure of a video interaction device provided in an embodiment of the present invention;

[0074] Figure 16 This is a schematic diagram of another video interaction device provided in an embodiment of the present invention;

[0075] Figure 17 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0076] To better understand the technical solutions of the embodiments of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0077] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0078] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0079] Real-time video interaction technologies such as live video streaming and video calls can acquire audio and video data shared from remote locations. However, during real-time video interaction, the display method of audio and video content is relatively fixed. For example, during a video call, the video content is always the image information of the video user, resulting in a poor user experience.

[0080] To address the aforementioned technical problems, embodiments of the present invention provide a video interaction method. This method, through 3D reconstruction, can display a 3D model of a target object within the video interaction frame, thereby further enhancing the interactive experience between the two parties and improving the user experience.

[0081] Figure 1 This is a schematic diagram of a video interaction scenario provided by an embodiment of the present invention. Figure 1 As shown, this video interaction scenario includes a first terminal side and a second terminal side. A video interaction connection is established between the first terminal side and the second terminal side. Optionally, the aforementioned video interaction connection can be a video call, video conference, live video streaming, etc. The first terminal side and the second terminal side can be smart terminal devices such as smartphones, tablets, smartwatches, and car central control screens; this embodiment of the invention is not limited to these.

[0082] like Figure 1 As shown, both the first and second terminal sides can acquire the audio and video content of the interacting parties. Figure 1 In this scenario, the first terminal displays a first video interaction screen, and the second terminal displays a second video interaction screen. The first and second video interaction screens are used to display video images from both parties involved in the video interaction.

[0083] Based on the real-time nature of video interactive connections, the first and second terminal sides can establish video interactive connections through the IP Multimedia Subsystem (IMS). Specifically, IMS is a subsystem superimposed on the existing packet domain in the Wide-band Code Division Multiple Access (WCDMA) network added in 3GPP Release 5. IMS uses the packet domain as its upper-layer control signaling and media transmission IMS channel. Specifically, the aforementioned upper-layer control signaling includes the Session Initiated Protocol (SIP) protocol. The SIP protocol, as the service control protocol of IMS, is used for session establishment and media negotiation between the first and second terminal sides. The aforementioned media transmission IMS channel includes audio and video channels. Based on the audio and video channels, audio and video stream data can be transmitted between the first and second terminal sides.

[0084] In some embodiments, an IP Multimedia Subsystem Data Channel (IMS DC) can be established between the first terminal and the second terminal. Optionally, the IMS DC can be used to transmit relevant data such as the 3D model of the target object. For example, in application scenarios such as live-streaming e-commerce and distance education, to more intuitively display the target object, the two parties in the video interaction can perform augmented reality processing on the video frame sequence of the target object contained in the video stream. The augmented target object is displayed in the video interaction screen in the form of a 3D model, providing convenience for user communication.

[0085] See Figure 2This is a flowchart illustrating a video interaction method provided in an embodiment of the present invention. In this method, a video interaction connection is established between a first terminal and a second terminal. The first terminal displays a first video interaction screen. The second terminal displays a second video interaction screen. In response to the aforementioned video interaction connection, the first terminal sends locally acquired video frames to the second terminal via an audio / video channel. Optionally, the video frames acquired by the first terminal can be video frames about a first target object. Optionally, the video frames about the first target object can be displayed in real-time on the second video interaction screen of the second terminal. In this embodiment of the present invention, the second terminal can also perform augmented reality processing on the video frames displayed on the second video interaction screen to construct a three-dimensional model of the first target object, so as to more intuitively and three-dimensionally display the first target object through the three-dimensional model.

[0086] like Figure 2 As shown, the processing steps of this method may include:

[0087] Step 101: The first terminal acquires a video frame sequence about the first target object. The video frame sequence includes video frames from various shooting angles of the first target object. Optionally, the first terminal can acquire the video frame sequence about the first target object in a video interaction scenario.

[0088] Step 102: The first terminal sends a video frame sequence about the first target object to the second terminal. Optionally, the first terminal can send the video frame sequence to the second terminal via an audio / video channel. Accordingly, the steps performed by the second terminal are described in steps 203 and 205.

[0089] Step 103: The second terminal receives the video frame sequence about the first target object sent by the first terminal. After receiving the video frame sequence about the first target object, the second terminal can display the video image about the first target object in real time on its second video interaction screen.

[0090] Step 104: The second terminal generates a 3D model of the first target object based on the video frame sequence of the first target object. Optionally, the second terminal not only displays video footage of the first target object in the second video interaction screen, but can also further generate a 3D model of the first target object based on the video frame sequence of the first target object.

[0091] Step 105: Display the 3D model of the first target object in the second video interaction screen on the second terminal side. Optionally, the video frame sequence of the first target object can be displayed in real time on the second terminal side in the manner of ordinary video interaction. Unlike this display method, the 3D model of the first target object generated on the second terminal side can be kept on display on the second terminal side at all times.

[0092] The video interaction method provided in this invention constructs a three-dimensional model of a first target object based on video frame sequences taken from various shooting angles. This three-dimensional model can be displayed in the video interaction screen, thereby improving the user experience during video interaction.

[0093] See Figure 3 This is a flowchart illustrating another video interaction method provided in an embodiment of the present invention. Figure 2 Unlike the methods shown, in this embodiment of the invention, after the first terminal side obtains a video frame sequence about the first target object, the first terminal side can construct a three-dimensional model based on the video frame sequence of the first target object. For example... Figure 3 As shown, the processing steps of this method include:

[0094] Step 201: The first terminal acquires a video frame sequence about the first target object. This video frame sequence includes video frames from various shooting angles of the first target object. Optionally, the first terminal can acquire the video frame sequence about the first target object within a video interaction scenario.

[0095] Step 202: The first terminal generates a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0096] Step 203: The first terminal sends the 3D model of the first target object to the second terminal. Optionally, the first terminal can send the 3D model of the first target object to the second terminal via IMS DC. The corresponding steps performed by the second terminal include steps 304 and 305 below.

[0097] Step 204: The second terminal receives the three-dimensional model of the first target object sent by the first terminal.

[0098] Step 205: The second terminal displays a 3D model of the first target object in the second video interaction screen.

[0099] In this embodiment of the invention, the first terminal acquires a video frame sequence about the first target object and constructs a three-dimensional model, and the second terminal can directly receive and display the three-dimensional model of the first target object in the video interaction screen.

[0100] In some embodiments, after the first terminal constructs a 3D model, it can also display the 3D model of the first target object in its own first video interaction screen. In this embodiment of the invention, both the first terminal and the second terminal display the 3D model of the first target object, which facilitates simultaneous display, explanation, or inquiry about the first target object by both parties.

[0101] In video interaction scenarios, to distinguish between ordinary video frames and video frames used to construct 3D models, the first video interaction screen on the first terminal side may include a first Augmented Reality (AR) function button. When the first terminal side detects the user's first trigger action on the first AR function button, the first terminal side activates the first AR mode. Optionally, the first terminal side determines the video frames captured in the first AR mode as video frames about the first target object.

[0102] In some embodiments, the second video interaction screen on the second terminal side may include a second AR function button. When the second terminal side detects a second trigger action by the user on the second AR function button, the second terminal side activates the second AR mode. The second terminal side determines the video frames received from the first terminal side in the second AR mode as video frames about the first target object. Optionally, the video frames about the first target object received in the second AR mode are video frames that can be used to construct a 3D model.

[0103] In some embodiments, the first AR mode described above can be an AR transmitting mode, i.e., used to transmit a sequence of video frames acquired about a first target object to a second terminal. The second AR mode described above can be an AR receiving mode, i.e., receiving a sequence of video frames used to construct a 3D model.

[0104] In some embodiments, the first AR function button can be a single button, and the first AR mode can be turned on or off by operating the single button. In some embodiments, the first AR function button can include at least a send-on sub-button and a send-off sub-button, wherein the send-on sub-button is used to turn on the AR sending mode and the send-off sub-button is used to turn off the AR sending mode.

[0105] In some embodiments, the second AR function button can be a single button, which can be used to turn the second AR mode on or off. In some embodiments, the second AR function button includes at least a receive-on sub-button and a receive-off sub-button, wherein the receive-on sub-button is used to turn on the AR receiving mode and the receive-off sub-button is used to turn off the AR receiving mode.

[0106] In some embodiments, the first terminal may simultaneously provide a first AR function button and a second AR function button to enable either the first AR mode or the second AR mode on the first terminal. Similarly, the second terminal may also simultaneously provide a first AR function button and a second AR function button to enable either the first AR mode or the second AR mode on the second terminal.

[0107] In some embodiments, the first AR function button and the second AR function button described above are buttons with the same function, both used to activate the AR transmission mode. Correspondingly, the first AR mode and the second AR mode described above are AR transmission modes. Optionally, the terminal device that activates the AR transmission mode captures and transmits a sequence of video frames about the first target object. Correspondingly, the terminal device that activates the AR transmission mode can notify the other device to receive the video frame sequence via instruction.

[0108] In some embodiments, IMS DC is supported between the first terminal and the second terminal. After acquiring a video frame sequence of the first target object, the first terminal decides which terminal device should perform the construction of the 3D model based on the device capabilities of both the first and second terminals. In some embodiments, if the device capability of the first terminal is less than or equal to that of the second terminal, the first terminal sends the video frame sequence of the first target object to the second terminal, so that the second terminal can construct the 3D model based on the video frame sequence of the first target object. If the device capability of the first terminal is greater than that of the second terminal, the first terminal constructs the 3D model based on the video frame sequence of the first target object.

[0109] Optionally, the device capabilities of the first terminal and the second terminal are determined based on the System-on-Chip (SOC) chip information of the first terminal and the second terminal, respectively. Specifically, the SOC chip information corresponding to the first terminal and the second terminal can be provided by the operator providing video interaction services to the first terminal and the second terminal. For example, the operator providing the video interaction service maintains a ranking chart of chip capabilities in the market in real time. The ranking chart includes SOC chip information for various types of terminal devices and their corresponding device capabilities.

[0110] In some embodiments, IMS DC is not supported between the first terminal and the second terminal. In this case, the first terminal sends a sequence of video frames about the first target object to the second terminal, which then constructs a 3D model based on the video frame sequence of the first target object.

[0111] The video interaction method of the present invention will be described in detail below with reference to specific embodiments.

[0112] See Figure 4 This is a flowchart illustrating another video interaction method provided by an embodiment of the present invention. In this embodiment, a video call is established between a first terminal and a second terminal, and IMSDC is not supported between the first terminal and the second terminal. Figure 4 As shown, the processing steps of this method may include:

[0113] Step 301: A video call connection is established between the first terminal side and the second terminal side, and IMS DC is not supported between the first terminal side and the second terminal side.

[0114] like Figure 5 As shown, the first terminal displays a first video interaction screen, and the second terminal displays a second video interaction screen. Both the first video interaction screen on the first terminal and the second video interaction screen on the second terminal include the local video screen and the peer video screen.

[0115] like Figure 5 As shown, the first video interaction screen on the first terminal side includes a first AR function button and a second AR function button. Similarly, the second video interaction screen on the second terminal side also includes a first AR function button and a second AR function button. The first AR function button is used to activate the AR sending mode. The second AR function button is used to activate the AR receiving mode. Optionally, such as... Figure 5 In this scenario, the first AR function button and the second AR function button on the first terminal side can be displayed in the local video screen on the first terminal side. Similarly, the first AR function button and the second AR function button on the second terminal side can also be displayed in the local video screen on the second terminal side.

[0116] During a video call between the first and second terminals, the audio and video captured by the first terminal are transmitted to the second terminal in real time, and the video feed from the first terminal is displayed in real time on the second video interaction screen of the second terminal. Similarly, the audio and video captured by the second terminal are transmitted to the first terminal in real time, and the video feed from the second terminal is displayed in real time on the first video interaction screen of the first terminal.

[0117] Step 302: When the first terminal detects a click action on the first AR function button in the first video interaction screen, the first terminal activates the AR sending mode. The video frames captured in AR sending mode are a sequence of video frames about the first target object.

[0118] In some embodiments, in AR sending mode, a user can hold the first terminal side and take a picture of the first target object. Optionally, the first terminal side can also provide shooting angle prompts to guide the user to take pictures of the first target object from multiple shooting angles.

[0119] After the first terminal side activates the AR sending mode, it can collect the user's voice information to instruct the second terminal side to activate the AR receiving mode.

[0120] Step 303: When the first terminal detects another click action on the first AR function button, it disables the AR transmission mode. After disabling the AR transmission mode, the video frame sequence collected by the first terminal is a normal video frame, no longer the video frame used to construct the 3D model. Optionally, after disabling the AR transmission mode, the first terminal can collect the user's voice information, which is used to instruct the second terminal to disable the AR receiving mode.

[0121] Step 304: The first terminal sends the video frame sequence captured in AR transmission mode to the second terminal. The video frame sequence captured by the first terminal in AR transmission mode is the video frame sequence about the first target object. Optionally, step 504 continues to be executed after AR transmission mode is enabled. Optionally, after AR transmission mode is switched on and off, the first terminal still sends the acquired video frame sequence to the second terminal; video frames after AR transmission mode is disabled are no longer used for AR 3D reconstruction.

[0122] The corresponding steps executed on the second terminal side include:

[0123] Step 305: When the second terminal detects a click action on the second AR function button in the second video interaction screen, the second terminal activates the AR receiving mode. Optionally, the user on the second terminal can activate the AR receiving mode based on voice information transmitted by the user on the first terminal.

[0124] Step 306: When the second terminal detects another click action on the second AR function button in the second video interaction screen, the second terminal disables the AR receiving mode. Optionally, the user on the second terminal can disable the AR receiving mode based on voice information transmitted by the user on the first terminal.

[0125] In step 307, the second terminal determines the video frame sequence acquired in AR receiving mode as a video frame sequence about the first target object. Optionally, step 507 continues to be executed after AR receiving mode is enabled. After AR receiving mode is disabled, the second terminal continues to receive audio and video information sent by the first terminal, and the video frames after AR receiving mode is disabled are no longer used for 3D reconstruction.

[0126] Step 308: The second terminal generates a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0127] Step 309: The second terminal displays the 3D model of the first target object in the second video interaction screen. Optionally, such as... Figure 5In this process, the 3D model of the first target object can be displayed in the local video frame of the second video interaction screen. Alternatively, the 3D model of the first target object can also be displayed in the remote video frame of the second video interaction screen; this invention is not limited to this.

[0128] In this embodiment of the invention, during a video call, the first terminal and the second terminal can determine the video frame sequence of the target object by enabling AR sending and AR receiving modes. The second terminal can then construct a three-dimensional model based on this video frame sequence. This three-dimensional model can be continuously displayed on the second video interaction screen of the second terminal, facilitating continuous and multi-angle display of the target object.

[0129] In some embodiments, Figure 4 and Figure 5 The method shown can also be applied to live video streaming scenarios. The first terminal can be the live streamer, and the second terminal can be the live stream receiver. The processing method in this embodiment can be further referred to the processing method provided in the above embodiments, and will not be repeated here. It should be noted that in this embodiment, the number of second terminals can be multiple, and is not limited here.

[0130] See Figure 6 This is a flowchart illustrating another video interaction method provided in an embodiment of the present invention. Figure 6 As shown, the processing steps of this method may include:

[0131] Step 401: A video call connection is established between the first terminal and the second terminal, and IMS DC is supported between the first terminal and the second terminal. After the video call connection is established between the first terminal and the second terminal, normal audio and video calls can be conducted between the first terminal and the second terminal.

[0132] like Figure 7 As shown, the first video interaction screen on the first terminal side includes a first AR function button. Similarly, the second video interaction screen on the second terminal side also includes a first AR function button. This first AR function button is used to activate the AR sending mode.

[0133] Step 402: When the first terminal detects a click action on the first AR function button in the first video interaction screen, the first terminal activates the AR sending mode.

[0134] The video frames of the first target object captured by the first terminal are sent to the second terminal through the audio and video channel and displayed in real time on the second video interactive screen of the second terminal.

[0135] Step 403: After the first terminal enables AR transmission mode, determine whether the device capability of the first terminal is greater than that of the second terminal. If yes, proceed to step 404. If no, proceed to step 408.

[0136] Step 404: When the first terminal detects a click action on the AR function button in the first video interaction screen again, the first terminal turns off the AR sending mode.

[0137] Step 405: The first terminal side determines the video frames captured in AR transmission mode as a video frame sequence about the first target object.

[0138] Step 406: The first terminal side constructs a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0139] Step 407: The first terminal sends the 3D model of the first target object to the second terminal via IMS DC.

[0140] Correspondingly, the second terminal can receive the 3D model of the first target object sent by the first terminal. The 3D model of the first target object is then displayed on the second video interaction screen of the second terminal. Optionally, such as... Figure 7 In this process, the 3D model of the first target object can be displayed on the counterpart video screen of the second video interaction screen. Alternatively, the 3D model of the first target object can also be displayed on the local video screen of the second video interaction screen; this invention is not limited in this respect.

[0141] Step 408: The first terminal sends a first instruction to the second terminal via IMS DC. The first instruction is used to instruct the second terminal to receive a video frame sequence about the first target object.

[0142] Step 409: When the first terminal detects a click action on the AR function button in the first video interaction screen again, the first terminal turns off the AR transmission mode and sends a second instruction to the second terminal via IMS DC. The second instruction is used to instruct the second terminal to stop receiving video frame sequences about the first target object.

[0143] Step 410: The second terminal side determines the video frames received between the first instruction and the second instruction as a video frame sequence about the first target object.

[0144] Step 411: The second terminal constructs a 3D model of the first target object based on the video frame sequence of the first target object. Optionally, after the second terminal constructs the 3D model of the first target object, it can display the 3D model of the first target object in the local video frame of the second video interaction screen. Alternatively, the 3D model of the first target object can also be displayed in the remote video frame of the second video interaction screen; this invention is not limited thereto.

[0145] In some embodiments, after the second terminal displays the 3D model of the first target object in the second video interaction screen, the method further includes: when a second operation command is detected acting on the 3D model of the first target object in the second video interaction screen, adjusting the display effect of the 3D model of the first target object according to the second operation command. In this way, the target object can be displayed from multiple angles and in a three-dimensional manner, further enhancing the user's experience in video interaction.

[0146] In some embodiments, after the second terminal side constructs a three-dimensional model of the first target object, it can send the three-dimensional model to the first terminal side so that the three-dimensional model of the first target object can also be displayed on the first terminal side.

[0147] In some embodiments, after the first terminal constructs a 3D model of the first target object, it not only sends it to the second terminal for display, but also displays the 3D model of the first target object in the first video interaction screen of the first terminal.

[0148] Correspondingly, when both the first terminal and the second terminal display a 3D model of the first target object, such as Figure 8 As shown, the method further includes:

[0149] Step 501: When the first terminal detects a preset operation applied to the 3D model in the first video interaction screen, it generates a first operation command. Optionally, such as... Figure 9 In this context, the aforementioned preset operation can be a long press and drag operation on the aforementioned 3D model, used to move or rotate the aforementioned 3D model.

[0150] Step 502: The first terminal responds to the first operation command and adjusts the display effect of the 3D model in the first video interaction screen. Based on the above preset operation, the first operation command can be an operation command to control the rotation of the 3D model. For example... Figure 9 In the process, the first terminal responds to the first operation command and rotates the 3D model displayed on the local video screen.

[0151] Step 503: The first terminal sends the first operation instruction to the second terminal.

[0152] In one possible implementation, the first operation command can be transmitted to the second terminal via the IMS DC.

[0153] Step 504: The second terminal responds to the first operation command and adjusts the display effect of the 3D model in the second video interaction screen. For example, ... Figure 9 As shown, the second terminal rotates the three-dimensional model displayed on the second terminal according to the first operation command sent by the first terminal, so as to control the three-dimensional model in the second terminal to present the same display effect as the three-dimensional model on the first terminal.

[0154] Similarly, the second terminal can also adjust the display effect of the 3D model on the first and second terminals according to preset operations applied to the 3D model in the second video interaction screen. In one possible implementation, the processing steps of this embodiment may further include:

[0155] Step 505: When the second terminal detects a preset operation applied to the 3D model in the second video interaction screen, it generates a second operation command.

[0156] Step 506: The second terminal responds to the second operation command and adjusts the display effect of the 3D model in the second video interaction screen.

[0157] Step 507: The second terminal sends the second operation instruction to the first terminal.

[0158] Step 508: The first terminal responds to the second operation command and adjusts the display effect of the 3D model in the first video interaction screen.

[0159] Specifically, in this embodiment, adjusting the display effect of the three-dimensional model includes, but is not limited to, rotating, moving, enlarging, and shrinking the three-dimensional model in the video interaction screen. This embodiment of the invention does not impose any limitations.

[0160] In applications such as video calls and live video streaming, this method allows users to explain or ask questions about the target object, making it easier for users to understand the target object more intuitively and improving the user experience.

[0161] Combination Figures 2 to 9 In some embodiments of the method shown, both the first terminal and the second terminal can be used to construct a 3D model of the first target object. Specifically, see [link to relevant documentation]. Figure 10 The above is a flowchart of a three-dimensional model construction method provided in an embodiment of the present invention.

[0162] like Figure 10 As shown, the processing steps of this method include:

[0163] Step 601: Preprocess the video frame sequence of the first target object to obtain the first target video frame sequence.

[0164] The aforementioned preprocessing of the video frame sequence of the first target object includes filtering out video frames with excessive repetition. It is understood that during the continuous acquisition of images of the first target object from various angles by the camera, images with excessive repetition will be generated. These highly repetitive images do not provide multi-view image information. If all of them are used for subsequent 3D reconstruction, it will introduce excessive redundant information for feature detection and point cloud reconstruction, significantly consuming the computing power of the mobile device without significantly improving reconstruction accuracy.

[0165] In addition, considering the instability of images captured by handheld devices, the preprocessing of the video frame sequence of the first target object also includes filtering out blurry or noisy video frames in the video frame sequence, thereby avoiding the introduction of errors in the subsequent model construction process and affecting the reconstruction accuracy of the 3D model.

[0166] Step 602: Extract features from each video frame in the first target video frame sequence using the Speeded-Up Robust Features (SURF) algorithm to obtain feature point descriptors for each video frame in the first target video frame sequence.

[0167] Specifically, compared to the Scale-invariant Feature Transform (SIFT) algorithm, the SURF algorithm is faster in computation and can also improve rotation robustness and scale transformation robustness to a certain extent.

[0168] Step 603: Based on the Euclidean distance between the feature point descriptors of each video frame, perform feature matching on the feature point descriptors of each video frame to obtain several feature point pairs.

[0169] Step 604: Construct a sparse point cloud of the first target object based on the relative positional relationship of each video frame in the first target video frame sequence and the feature point pairs.

[0170] Specifically, the above-mentioned construction of the sparse point cloud of the first target object also includes: using bundle adjustment to eliminate accumulated errors.

[0171] Step 605: Determine the depth value of the feature point descriptor of each video frame based on the sparse point cloud of the first target object.

[0172] Step 606: Construct a dense point cloud of the first target object based on the depth values ​​of the feature point descriptors of each video frame.

[0173] Step 607: Perform mesh reconstruction and simplification smoothing on the dense point cloud to obtain a three-dimensional model of the first target object.

[0174] Based on the three-dimensional model construction method provided in this embodiment, a three-dimensional model of the first target object is constructed according to the video frame sequence of the first target object, thereby enabling the three-dimensional reconstruction and preview of the first target object and improving the user experience.

[0175] Figure 11 This is a schematic diagram of a video interaction device provided in an embodiment of the present invention. The video interaction device can be applied to a first terminal side. Figure 11 As shown, the aforementioned video interaction device may include:

[0176] The first acquisition module 71 is specifically used to acquire a video frame sequence of a first target object, the video frame sequence including video frames of the first target object from various shooting angles.

[0177] The first transmission module 72 is specifically used to send a video frame sequence of the first target object to the second terminal side. The video frame sequence of the first target object is used to generate a three-dimensional model of the first target object on the second terminal side.

[0178] Figure 12 This is a schematic diagram of another video interaction device provided in an embodiment of the present invention, which can be applied to a first terminal side. For example... Figure 12 As shown, the video interaction device may further include: a first acquisition module 71, a first generation module 73, and a first transmission module 72.

[0179] The first acquisition module 71 is specifically used to acquire a video frame sequence of a first target object, the video frame sequence including video frames of the first target object from various shooting angles.

[0180] The first generation module 73 is specifically used to generate a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0181] The first transmission module 72 is also used to send the three-dimensional model of the first target object to the second terminal side.

[0182] Figure 13 This is a schematic diagram of a video interaction device provided in an embodiment of the present invention. The video interaction device can be applied to a second terminal side. Figure 13 As shown, the aforementioned video interaction device may include:

[0183] The second transmission module 81 is specifically used to receive the video frame sequence of the first target object sent by the first terminal side.

[0184] The second generation module 82 is specifically used to generate a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0185] The second display module 83 is specifically used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side.

[0186] Figure 14 This is a schematic diagram of another video interaction device provided in an embodiment of the present invention, which can be applied to a second terminal side. For example... Figure 14 As shown, the above-mentioned video interaction device may include: a second transmission module 81 and a second display module 83.

[0187] The second transmission module 11 is specifically used to receive the three-dimensional model of the first target object sent by the first terminal side.

[0188] The second display module 83 is specifically used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side.

[0189] Figure 15 This is a schematic diagram of the structure of a video interaction device provided in an embodiment of the present invention. Figure 15 As shown, the video interaction device may include a first acquisition module 71, a first transmission module 72, a second transmission module 81, a second generation module 82, and a second display module 83.

[0190] In this embodiment, the first acquisition module 71 is specifically used to acquire a video frame sequence about the first target object, the video frame sequence including video frames of the first target object from various shooting angles.

[0191] The first transmission module 72 is used to send the video frame sequence of the first target object to the second terminal side.

[0192] The second transmission module 81 is specifically used to receive the video frame sequence of the first target object sent by the first terminal side.

[0193] The second generation module 82 is specifically used to construct a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0194] The second display module 83 is specifically used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side.

[0195] In addition, such as Figure 15 As shown, the first terminal and the second terminal establish a video interaction connection through the first transmission module 72 and the second transmission module 81. This video interaction connection can be a video call connection or a live video connection, and is not limited here.

[0196] Figure 16 This is a schematic diagram of another video interaction device provided in an embodiment of the present invention. Figure 16As shown, the video interaction device may include: a first acquisition module 71, a first generation module 73, a first transmission module 72, a second transmission module 81, and a second display module 83.

[0197] The first acquisition module 71 is specifically used to acquire a video frame sequence about a first target object, the video frame sequence including video frames of the first target object from various shooting angles.

[0198] The first generation module 73 is specifically used to construct a three-dimensional model of the first target object based on the video frame sequence of the first target object.

[0199] The first transmission module 72 is used to send the three-dimensional model of the first target object to the second terminal side.

[0200] The second transmission module 81 is specifically used to receive the three-dimensional model of the first target object sent by the first terminal side.

[0201] The second display module 83 is specifically used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side.

[0202] The video interaction method provided in this invention supports the reconstruction and preview of a 3D model based on video frame data of a specific object during video interaction. Simultaneously, the 3D model can be transmitted along with the audio and video data to the other terminal for display, facilitating user communication.

[0203] Figure 17 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 17 The electronic device shown can be implemented as the first terminal side described above, or it can also be implemented as the second terminal side described above. For example... Figure 17 As shown, the electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors 910, memory 930, and communication interface 920, and a communication bus 940 connecting different system components (including processor 910, communication interface 920, and memory 930).

[0204] The communication bus 940 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0205] Electronic devices typically include a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, and removable and non-removable media.

[0206] Memory 930 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Memory 930 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0207] A program / utility having a set (at least one) of program modules can be stored in memory 930. Such program modules include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules typically perform the functions and / or methods described in the embodiments of the present invention.

[0208] Processor 910 executes various functional applications and data processing by running programs stored in memory 930, such as implementing embodiments of the present invention. Figure 2 , Figures 4-5 or Figure 3 , Figures 6-9 The illustrated embodiment provides a video interaction method.

[0209] This invention provides a non-transitory computer-readable storage medium that stores computer instructions, which cause the computer to execute embodiments of this invention. Figure 2 , Figures 4-5 or Figure 3 , Figures 6-9 The illustrated embodiment provides a video interaction method.

[0210] The aforementioned non-transitory computer-readable storage medium may be any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in connection with an instruction execution system, apparatus, or device.

[0211] The foregoing has described specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0212] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of embodiments of the present invention, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0213] In the several embodiments provided in this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0214] Furthermore, in the various embodiments of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0215] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A video interaction method, characterized in that, The method is applied to a first terminal side, where a video interactive connection is established between the first terminal side and a second terminal side. The method includes: Obtain a video frame sequence of a first target object, the video frame sequence containing video frames of the first target object from various shooting angles; send the video frame sequence to a second terminal, the video frame sequence being used to generate a 3D model of the first target object on the second terminal; or, Obtain the video frame sequence of the first target object, generate a three-dimensional model of the first target object based on the video frame sequence, and send the three-dimensional model of the first target object to the second terminal side; Sending the video frame sequence to the second terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the video frame sequence is sent to the second terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or, If the IP Multimedia Subsystem data channel is not supported between the first terminal side and the second terminal side, the video frame sequence is obtained and then sent to the second terminal side. The step of generating a 3D model of the first target object based on the video frame sequence includes: In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, a three-dimensional model of the first target object is generated based on the video frame sequence, and the three-dimensional model of the first target object is sent to the second terminal through the IP Multimedia Subsystem data channel.

2. The method according to claim 1, characterized in that, The acquisition of the video frame sequence about the first target object includes: The first terminal displays a first video interaction screen, which includes a first augmented reality (AR) function button. After detecting the first trigger action of the first AR function button, the first AR mode is activated, wherein the video frames captured in the first AR mode are a sequence of video frames about the first target object.

3. The method according to claim 2, characterized in that, After activating the first AR mode, the method further includes: Send a first instruction to the second terminal side, the first instruction being used to instruct the second terminal side to determine a video frame sequence about the first target object from the received video frames.

4. The method according to claim 1, characterized in that, After generating the 3D model of the first target object based on the video frame sequence, the method further includes: A three-dimensional model of the first target object is displayed in the first video interactive screen on the first terminal side.

5. The method according to claim 4, characterized in that, The method further includes: When a first operation command is detected that acts on the three-dimensional model of the first target object in the first video interaction screen, the display effect of the three-dimensional model of the first target object is adjusted according to the first operation command. The first operation instruction is sent to the second terminal side, and the second terminal side displays a second video interaction screen. The second video interaction screen displays a three-dimensional model of the first target object. The first operation instruction is also used by the second terminal side to adjust the display effect of the three-dimensional model in the second video interaction screen.

6. The method according to claim 4, characterized in that, The method further includes: Receive a second operation command sent by the second terminal side, the second operation command acting on the three-dimensional model in the second video interaction screen of the second terminal side; Adjust the display effect of the three-dimensional model of the first target object in the first video interaction screen according to the second operation instruction.

7. A video interaction method, characterized in that, The method is applied to a second terminal side, where a video interactive connection is established between the second terminal side and the first terminal side. The method includes: Receive a sequence of video frames about a first target object sent by a first terminal, generate a three-dimensional model of the first target object based on the video frame sequence, and display the three-dimensional model of the first target object in a second video interactive screen on a second terminal; or, Receive a 3D model of a first target object sent from the first terminal side; display the 3D model of the first target object in the second video interactive screen on the second terminal side; The method of receiving the video frame sequence about the first target object sent by the first terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the system receives a video frame sequence about the first target object sent by the first terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or... In response to the lack of support for an IP Multimedia Subsystem data channel between the second terminal and the first terminal, the system receives a sequence of video frames about the first target object sent by the first terminal. The three-dimensional model of the first target object received from the first terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, the device receives a three-dimensional model of the first target object sent by the first terminal through the IP Multimedia Subsystem data channel.

8. The method according to claim 7, characterized in that, The video frame sequence about the first target object sent by the first terminal includes: The second video interaction screen on the second terminal side includes a second AR function button; after detecting a second trigger action on the second AR function button, a second AR mode is activated, wherein the video frames received from the first terminal side in the second AR mode are a sequence of video frames about the first target object; or... After receiving the first instruction sent by the first terminal, a video frame sequence about the first target object is determined from the video frames sent by the first terminal according to the first instruction.

9. The method according to claim 7, characterized in that, After receiving the three-dimensional model of the first target object sent by the first terminal side, the method further includes: Receive the first operation instruction sent by the first terminal side; Adjust the display effect of the three-dimensional model of the first target object in the second video interaction screen according to the first operation instruction.

10. The method according to claim 9, characterized in that, After receiving the three-dimensional model of the first target object sent by the first terminal side, the method further includes: When a second operation command is detected that acts on the three-dimensional model of the first target object in the second video interaction screen, the second operation command is sent to the first terminal side. The second operation command is also used by the first terminal side to adjust the display effect of the three-dimensional model of the first target object in the first video interaction screen displayed on the first terminal side.

11. The method according to claim 7, characterized in that, The method further includes: When a second operation command is detected that acts on the 3D model of the first target object in the second video interaction screen, the display effect of the 3D model of the first target object is adjusted according to the second operation command.

12. A video interactive device, characterized in that, Applied to the first terminal side, including: The first acquisition module is used to acquire a video frame sequence of a first target object, wherein the video frame sequence contains video frames of the first target object from various shooting angles. The first transmission module is used to send a video frame sequence of the first target object to the second terminal side, and the video frame sequence of the first target object is used to generate a three-dimensional model of the first target object on the second terminal side. or, The video interaction device further includes: a first generation module, configured to generate a three-dimensional model of the first target object based on a video frame sequence of the first target object; and a first transmission module, configured to send the three-dimensional model of the first target object to the second terminal side. Sending the video frame sequence to the second terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the video frame sequence is sent to the second terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or, If the IP Multimedia Subsystem data channel is not supported between the first terminal side and the second terminal side, the video frame sequence is obtained and then sent to the second terminal side. The step of generating a 3D model of the first target object based on the video frame sequence includes: In response to the support of an IP Multimedia Subsystem data channel between the first terminal and the second terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, a three-dimensional model of the first target object is generated based on the video frame sequence, and the three-dimensional model of the first target object is sent to the second terminal through the IP Multimedia Subsystem data channel.

13. A video interactive device, characterized in that, Applied to the second terminal side, including: The second transmission module is used to receive a video frame sequence of the first target object sent by the first terminal side. The second generation module is used to generate a three-dimensional model of the first target object based on the video frame sequence of the first target object; The second display module is used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side; or, The second transmission module is used to receive the three-dimensional model of the first target object sent by the first terminal side; the second display module is used to display the three-dimensional model of the first target object in the second video interactive screen on the second terminal side. The method of receiving the video frame sequence about the first target object sent by the first terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the device capability of the first terminal being less than or equal to the device capability of the second terminal, the system receives a video frame sequence about the first target object sent by the first terminal. The device capabilities of the first terminal and the second terminal are determined based on the system-on-chip (SoC) chip information of the first terminal and the second terminal, respectively; or... In response to the lack of support for an IP Multimedia Subsystem data channel between the second terminal and the first terminal, the system receives a sequence of video frames about the first target object sent by the first terminal. The three-dimensional model of the first target object received from the first terminal includes: In response to the support of an IP Multimedia Subsystem data channel between the second terminal and the first terminal, and the fact that the device capability of the first terminal is greater than that of the second terminal, the device receives a three-dimensional model of the first target object sent by the first terminal through the IP Multimedia Subsystem data channel.

14. An electronic device, characterized in that, include: At least one processor; as well as At least one memory communicatively connected to the processor, wherein: The memory stores program instructions that can be executed by the processor, which can invoke the program instructions to perform the method as described in any one of claims 1 to 6 or 7-11.

15. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the method as described in any one of claims 1 to 6 or 7-11.

Citation Information

Patent Citations

  • Information processing method and device, electronic equipment and medium

    CN110336973A

  • AR interaction method and device of equipment, electronic equipment and storage medium

    CN112650422A

  • Method, device and equipment for playing real-time video stream

    CN115883814A