Video call establishment method and device, equipment, medium and program product
By employing holographic communication identifiers and a video stream negotiation mechanism with multiple transmission channels in the IMS network, the problem that single-camera video calls cannot meet the needs of multi-view holography is solved, and high-resolution synchronous transmission and immersive experience of multi-view video streams are achieved.
Patent Information
- Application Number
- CN202511710742.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-17
AI Technical Summary
Existing video call technology based on IMS networks cannot meet users' needs for immersive multi-view holographic video calls. The single-camera acquisition mode only provides 2D video images from a single perspective, which cannot achieve an immersive holographic call experience with multi-angle motion parallax.
Through the IMS network, the calling terminal sends a call request to the called terminal, carrying a holographic communication identifier and transmission channel information of multiple video streams. The called terminal generates a response message according to its capabilities. The calling terminal establishes a holographic or single-channel video call based on the response message, uses multiple transmission channels to transmit multi-view video streams, and uses a time-space stamping mechanism to process video frames synchronously.
It enables multi-view synchronous transmission of holographic video calls under the IMS network, avoiding the degradation of picture quality caused by the resolution limit, improving the clarity and immersion of video calls, and adapting to the stability and availability of different network environments.
Smart Images

Figure CN121547549A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication technology, and in particular to a method and apparatus for establishing a video call, a device, a medium and a program product. BACKGROUND
[0002] Holographic communication is a whole communication process covering data collection, coding, transmission, rendering and display, and is applied to high-immersion and multi-dimensional interaction scenarios. It includes the whole end-to-end process from data collection to multi-dimensional sensory data restoration, and is a new type of communication service form with high-immersion and high-naturalness interaction characteristics and motion parallax effect.
[0003] In the current field of telecommunications operator communication technology, communication application scenarios based on IMS networks are constantly expanding, and currently mainly cover voice calls, video calls and some basic data transmission services. Among them, audio and video call services rely on SIP protocol as the core for session negotiation. In this process, the two parties agree on parameters such as audio and video coding format, resolution and frame rate, thereby ensuring stable transmission of audio and video streams between different terminals, and ultimately realizing face-to-face visual communication.
[0004] Existing video call technology based on IMS networks can only use a single camera of a terminal device for collection. However, with the rise of vehicle cockpits and multi-camera intelligent terminals, the single-camera collection mode can only provide a single perspective 2D video screen, which cannot meet the needs of users for immersive multi-perspective holographic video calls, and cannot provide immersive holographic call experience with multi-angle motion parallax. SUMMARY
[0005] The embodiments of the present application provide a method and apparatus for establishing a video call, a device, a medium and a program product to solve the problem that existing video call technology based on IMS networks cannot meet the needs of users for immersive multi-perspective holographic video calls.
[0006] In a first aspect, the embodiments of the present application provide a method for establishing a video call, applied to a calling terminal in an IMS network, and the method comprises: sending a call request to a called terminal through an IMS network, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a multi-path video stream of the called terminal; receiving a response message sent by the called terminal through the IMS network, wherein the response message is generated by the called terminal based on the call request; establishing a video call with the called terminal based on the response message.
[0007] In a second aspect, the application provides a method for establishing a video call, applied to a called terminal in an IMS network, and the method comprises the following steps: receiving, through the IMS network, a call request sent by a calling terminal, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; generating a response message based on the call request; sending, through the IMS network, the response message to the calling terminal, wherein the response message is used to instruct the calling terminal to establish a video call with the called terminal.
[0008] In a third aspect, the application provides an apparatus for establishing a video call, deployed in a calling terminal in an IMS network, and the apparatus comprises the following modules: a sending module, configured to send, through the IMS network, a call request to a called terminal, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; a receiving module, configured to receive, through the IMS network, a response message sent by the called terminal, wherein the response message is generated by the called terminal based on the call request; an establishing module, configured to establish a video call with the called terminal based on the response message.
[0009] In a fourth aspect, the application provides an apparatus for establishing a video call, deployed in a called terminal in an IMS network, and the apparatus comprises the following modules: a receiving module, configured to receive, through the IMS network, a call request sent by a calling terminal, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; a generating module, configured to generate a response message based on the call request; a sending module, configured to send, through the IMS network, the response message to the calling terminal, wherein the response message is used to instruct the calling terminal to establish a video call with the called terminal.
[0010] In a fifth aspect, the application provides an electronic device, comprising a processor and a memory storing a computer program, wherein the processor implements the steps of the method for establishing a video call according to the first aspect or the second aspect when executing the program.
[0011] In a sixth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the steps of the method for establishing a video call according to the first aspect or the second aspect.
[0012] In a seventh aspect, an embodiment of the present application provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the method for establishing a video call according to the first aspect or the second aspect.
[0013] The present application provides a method for establishing a video call in an IMS network. A calling terminal sends a call request to a called terminal through an IMS network, the call request carrying a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; the called terminal receives a response message sent by the called terminal through the IMS network, the response message being generated by the called terminal based on the call request; and finally, the calling terminal establishes a video call with the called terminal based on a type of the response message. By implementing the method of the present application, the calling terminal first initiates a call request for establishing a holographic video call, and the called terminal generates a response message based on its own capability information, and the type of the response message determines the type of the video call to be finally established. In this way, the terminal can not only perform holographic video call in the IMS network, but also perform single-channel video call in the IMS network, thereby greatly improving the video call capability of the terminal in the IMS network and overcoming the problem in the related art that the terminal in the IMS network can only collect a single-channel video stream using a single camera, which cannot meet the user's demand for immersive multi-view holographic video call. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0015] Figure 1 is a flowchart of a method for establishing a video call according to an embodiment of the present application; Figure 2 is a flowchart of a method for establishing a holographic video call between a calling terminal and a called terminal according to an embodiment of the present application; Figure 3 is a flowchart of a method for establishing a single-channel video call between a calling terminal and a called terminal according to an embodiment of the present application; Figure 4is a flowchart of a process of updating information of a transmission channel, as illustrated in an embodiment of the present application; Figure 5 is a flowchart of another method of establishing a video call, as illustrated in an embodiment of the present application; Figure 6 is a schematic diagram of a process of completing missing video frames, as illustrated in an embodiment of the present application; Figure 7 is a structural block diagram of an apparatus for establishing a video call, as illustrated in an embodiment of the present application; Figure 8 is a structural block diagram of another apparatus for establishing a video call, as illustrated in an embodiment of the present application; Figure 9 is a schematic diagram of a physical structure of an electronic device, as illustrated in an embodiment of the present application. DETAILED DESCRIPTION
[0016] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be described below in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0017] To solve the problems in the prior art, the present application provides a method of establishing a video call, which is applicable to a video call under an IMS network. The execution subject of the method is a calling terminal. Figure 1 is a flowchart of a method of establishing a video call, as illustrated in an embodiment of the present application. Referring to Figure 1 , the method of the present application can include: Step S101, sending a call request to a called terminal through an IMS network, the call request carrying a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal.
[0018] The IMS (IP Multimedia Subsystem) is a standardized network architecture for providing multimedia communication services over an IP network. It is independent of access technologies (such as 5G, LTE, or Wi-Fi) and uses the Session Initiation Protocol (SIP) as the core to initiate, control, and negotiate voice calls, video calls, and data transmission services.
[0019] In this embodiment, when the calling terminal needs to make a video call with the called terminal, a call request is sent to the called terminal through an IMS network. The call request carries a first holographic communication identifier and information of a plurality of first transmission channels.
[0020] The first holographic communication identifier is used to inform the called terminal that the calling terminal has holographic communication capability and that the current call request aims to initiate a holographic video call. The information of the plurality of first transmission channels is determined by the calling terminal according to its own capability information, and is used to inform the called terminal to transmit a plurality of video streams to the calling terminal according to the information of the plurality of first transmission channels. The plurality of video streams are video streams collected by the called terminal through a plurality of different cameras, and one camera is used to collect a video stream at one angle.
[0021] In this embodiment, the calling terminal can determine the number of the plurality of first transmission channels according to the number of its own cameras.
[0022] In step S102, a response message sent by the called terminal is received through the IMS network. The response message is generated by the called terminal based on the call request.
[0023] In this embodiment, after receiving the response information, the called terminal determines which response to generate according to the content in the response information and the capability information of the called terminal. If the called terminal has holographic communication capability, a first response message is generated and sent to the calling terminal. If the called terminal does not have holographic communication capability, a second response message is generated and sent to the calling terminal.
[0024] In step S103, a video call between the calling terminal and the called terminal is established based on the response message.
[0025] In this embodiment, the calling terminal establishes a video call with the called terminal according to the type of the response message. If the type of the response message is the first response message, a holographic video call is established with the called terminal. If the type of the response message is the second response message, a single-channel video call is established with the called terminal.
[0026] The application provides a call establishment scheme for video call in an IMS network. A calling terminal sends a call request to a called terminal through the IMS network, the call request carrying a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; the called terminal receives a response message sent by the called terminal through the IMS network, the response message being generated by the called terminal based on the call request; and finally, the calling terminal establishes a video call with the called terminal based on the type of the response message. By implementing the method of the application, the calling terminal first initiates a call request for establishing a holographic video call, and the called terminal generates a response message based on its own capability information, and the type of the response message determines the type of the video call to be finally established. In this way, the terminal can not only perform holographic video call in the IMS network, but also perform single-channel video call in the IMS network, thereby greatly improving the video call capability of the terminal in the IMS network and overcoming the problem in the related art that the terminal in the IMS network can only use a single camera to collect a single-channel video stream, which cannot meet the user's demand for immersive multi-view holographic video call.
[0027] In an existing holographic call technology in a non-IMS network, in order to realize synchronous transmission of multi-view video streams, the multi-view video streams are combined into a grid picture for synchronous transmission. Specifically, the video streams from different cameras are spliced to form a grid picture (for example, a 9-grid picture), and then the integrated grid picture is synchronously transmitted through a real-time transport protocol (RTP) or real-time transport control protocol (RTCP) video stream. However, this method has obvious limitations: since there is an upper limit to the supported resolution of the RTP / RTCP video stream in the transmission process, as the number of cameras at the collection end increases, the resolution of the single-view video stream represented by each small grid in the grid picture will gradually decrease after splicing. For example, when the number of cameras increases from 2 to 5 or even more, the proportion of each view in the grid picture becomes smaller, resulting in loss of picture details and reduction of clarity when the final synthesized video picture presents a holographic effect, which affects the user's video call experience. To solve this problem, in an embodiment, step S103 can include: In the case where the response message carries a second holographic communication identifier indicating that the called terminal has holographic communication capability, and information of a plurality of second transmission channels for receiving a plurality of video streams of the calling terminal, the holographic video call with the called terminal is established; wherein the information of the plurality of second transmission channels is determined based on the information of the plurality of first transmission channels.
[0028] In the embodiment, after receiving the response message, the calling terminal specifically analyzes the content in the response message: if the response message contains both the second holographic communication identifier indicating that the called terminal has holographic communication capability and the information of the multiple second transmission channels for receiving the multiple video streams of the calling terminal, it indicates that the type of the response message is the first response message. Then, the calling terminal directly establishes a holographic video call with the called terminal.
[0029] In the embodiment, the calling terminal carries the information of the multiple first transmission channels suitable for transmitting the multiple video streams obtained by its own analysis when initiating the call request. When receiving the call request, if the called terminal confirms that it has holographic communication capability, it will agree to the content in the call request, and then generate a first response message and carry the information of the multiple second transmission channels suitable for transmitting the multiple video streams obtained by its own analysis in the first response message. In this way, when the calling terminal determines that the type of the response message is the first response message, it considers that the negotiation has been completed, and therefore can directly establish a holographic video call with the called terminal.
[0030] In the embodiment, instead of combining the multiple video streams into a mosaic and then transmitting, the multiple transmission channels obtained through negotiation are used to transmit the multiple video streams, so that the problem of reduced resolution of single-view video stream in a multi-camera scene caused by the upper limit of RTP / RTCP video stream resolution can be avoided. Since each video stream can be transmitted through a separate RTP / RTCP channel in the embodiment, the video stream of each view can be transmitted at a high resolution, avoiding the loss of picture quality caused by channel multiplexing, and significantly improving the picture clarity and detail performance of the holographic video call, providing users with a better visual experience.
[0031] In combination with the above embodiments, in an implementation, after step S103, the method of the present application can further include: obtaining the multiple video streams collected by the camera of the calling terminal, and the information of each video frame corresponding to the timestamp and the camera in each video stream; transmitting the multiple video streams, the timestamp corresponding to each video frame, and the information of the camera to the called terminal according to the information of the multiple second transmission channels.
[0032] In the embodiment, the calling terminal can transmit the multiple video streams to the called terminal after establishing the holographic video call with the called terminal. A uniform timestamp generation module is arranged in the calling terminal. When the multiple video streams of multiple perspectives are collected by the multiple cameras of the calling terminal, the uniform timestamp generation module marks the same timestamp for the different video frames collected at the same time by using a uniform clock source. Then, the calling terminal transmits the multiple video streams marked with the timestamp and the information of the cameras of the video frames to the called terminal according to the information of the multiple second transmission channels. In this way, the called terminal can use the received uniform timestamp to accurately time-align the multiple video streams collected at the same time, ensure that the pictures of all perspectives are strictly synchronized, and prevent delay or asynchronization between different perspectives. Meanwhile, by using the information of the cameras, the called terminal can accurately identify the spatial source corresponding to each video stream, so as to sort, reconstruct and render the pictures in the correct spatial relationship, and finally realize a coherent, immersive and spatially accurate holographic multi-perspective call experience.
[0033] The information of the camera includes the identity of the camera and the position of the camera, and can be set according to actual needs.
[0034] In the embodiment, the multiple video streams, the timestamps corresponding to the video frames, and the information of the cameras are transmitted to the called terminal according to the information of the multiple second transmission channels, which can enhance the holographic video call experience between the calling terminal and the called terminal. In addition, the space-time stamp (timestamp + camera position / identity) mechanism not only records the time information of the video frames, but also integrates the position or identity of the camera, so that the called terminal can sort and synchronize the video frames in double dimensions according to the space-time stamp, can fully consider the spatial relationship and time difference of different camera perspectives, effectively eliminate the picture asynchronization problem caused by factors such as shooting angle and transmission delay, ensure that the multiple perspective videos are highly consistent in time and space, and finally achieve the effect of significantly enhancing the immersion of the holographic video call.
[0035] In combination with the above embodiments, in an implementation, step S101 includes: The calling request is sent to the called terminal by using the SIP protocol, wherein the first holographic communication identifier is located in the SIP header field of the calling request, and the information of the multiple first transmission channels is located in the SDP message body of the calling request.
[0036] In the IMS network, the SIP protocol is used as the signaling control layer and is responsible for the establishment and management of the session. The signaling message body of the SIP protocol uses the Session Description Protocol (SDP) protocol to describe and negotiate the specific media information.
[0037] The application extends the SIP header field, and adds the Contact field to carry the holographic communication capability identifier, i.e. declares the terminal has the holographic call capability.
[0038] The following will describe how the calling terminal establishes the holographic video call with the called terminal in the application by an embodiment. In the embodiment, both the calling terminal and the called terminal have the holographic communication capability. Figure 2 is a flow chart for establishing the holographic video call between the calling terminal and the called terminal according to the embodiment of the application. Referring to Figure 2 , the embodiment includes the following steps: Step 1: The calling terminal sends the call request, i.e. the INVITE message, to the called terminal. The INVITE message is sent by using the SIP protocol, and the Contact: gfree-holography-video field (as the first holographic communication identifier) is carried in the SIP header field, which is used to indicate that the calling terminal has the holographic communication capability. In addition, the information of the multiple first transmission channels determined by the calling terminal (port, coding format, resolution, etc.) is carried in the SDP message body.
[0039] An example of the SIP header field in the INVITE message is as follows: The Contact: gfree-holography-video field in the SIP header field indicates that the calling terminal has the holographic communication capability INVITE tel:+86187XXXX2222 SIP / 2.0 Allow:INVITE, ACK, OPTIONS, CANCEL, BYE, UPDATE, INFO, REFER, NOTIFY, MESSAGE, PRACK Contact: +g.3gpp.icsi-ref="urn%3Aurn-7%3A3gpp-service.ims.icsi.mmtel"; +sip.instance=" <urn:gsma:imei:35293609-715199-0>" gfree-holography-video” In this context, `INVITE tel:+86187XXXX2222` indicates an invitation to +86187XXXX2222 for a video call. `SIP / 2.0` indicates that this call is based on SIP protocol version 2.0. The `Allow` field lists the SIP signaling methods supported by the calling terminal. `Allow: INVITE, ACK, OPTIONS...` indicates that the calling terminal supports initiating or receiving these types of signaling operations (e.g., initiating an invitation, acknowledgment, cancellation, etc.), which represents the calling terminal's capabilities at the protocol interaction layer. The content in the `Contact` field is the first holographic communication identifier.
[0040] Here is an example of an SDP message body in an INVITE message: # Video Stream: Viewpoint 1 m=video 5000 RTP / AVP 126 # Specifies the load type 126 corresponding to H.265 encoding, and the clock frequency is 90000Hz. a=rtpmap:126 H265 / 90000 # Define additional parameters for payload type 126: a=fmtp:126 # Supports mainstream application scenarios (real-time communication); Level 4.0 supports 1080p resolution. profile-id=1;level-id=120; # H.265 global encoding features (VPS), resolution, chroma format, bit depth (SPS), and frame-level parameters (PPS). sprop-vps= <video parameter sets data>; sprop-sps= <video parameter sets data>; sprop-pps= <video parameter sets data> # frame rate a=framerate:30 # identify the negotiation holographic multi-view video call media channel, transmit multi-view media stream a=content:visualized-voice-call,gfree-holography-video / / video stream: view 2 m=video 5002 RTP / AVP 127 ... / / video stream: view 3 m=video 5004 RTP / AVP 128 ... / / video stream: view 9 m=video 5018 RTP / AVP 135 # specify that the payload type 135 corresponds to H.265 encoding, clock frequency 90000Hz a=rtpmap:135 H265 / 90000 # define additional parameters for payload type 135: a=fmtp:135 # support the main application scenario (real-time communication); Level 4.0 supports 1080 resolution profile-id=1;level-id=120; # global encoding characteristics (vps) of H265, resolution, chroma format, bit depth (sps), frame-level parameters (pps) sprop-vps= <video parameter sets data>; sprop-sps= <video parameter sets data>; sprop-pps= <video parameter sets data> # frame rate a=framerate:30 # identify the holographic multi-view video call media channel, transmit multi-view media stream a=content:visualized-voice-call,gfree-holography-video” In this example, the number after m=video represents the port number for receiving the video stream corresponding to the current view, for example, in m=video 5000, 5000 is the port number. RTP / AVP indicates that the media stream is transmitted using the Real-time Transport Protocol (RTP) and its Audio Video Profile (AVP). The m line is used to declare a transmission channel (for example, using the payload 126), and the number after m=video is used as the port number to receive the video stream of the transmission channel. The a line is used to describe the information of the transmission channel of the m line above it. In this embodiment, the H.265 encoding method is used, and the minimum supported resolution is 1080.
[0041] Step 2: The called terminal returns a 183 Session Progress response to the calling terminal. The response is sent using the SIP protocol, and the Contact: gfree-holography-video field (which serves as the second holographic communication identifier) is carried in the SIP header field, indicating that the called terminal has holographic communication capability. Secondly, the information of the multiple second transmission channels determined by the called terminal (port, encoding format, resolution, etc.) is carried in the SDP message body.
[0042] An example of a SIP header field of 183 Session Progress is as follows: "SIP / 2.0 183 Session Progress Allow: INVITE, ACK, OPTIONS, CANCEL, BYE, UPDATE, INFO, REFER, NOTIFY, MESSAGE, PRACK Contact: +g.3gpp.icsi-ref="urn%3Aurn-7%3A3gpp-service.ims.icsi.mmtel"; +sip.instance=" <urn:gsma:imei:35293609-715199-0> gfree-holography-video” Wherein, SIP / 2.0 indicates that this call is based on the SIP protocol version 2.0. Allow: INVITE, ACK, OPTIONS... indicates that the called terminal supports initiating or receiving these types of signaling operations, which is the capability of the called terminal at the protocol interaction level. The content in the Contact field is the second holographic communication identifier.
[0043] Step 3: After receiving the 183 Session Progress response, the calling terminal sends a SIP PRACK (Provisional Acknowledgement) to confirm the media parameters.
[0044] Step 4: After receiving the PRACK message, the called user terminal replies with a 200 OK response to the PRACK message.
[0045] Step 5: The called terminal starts ringing and sends a 180 Ringing response.
[0046] Step 6: After the called terminal answers the holographic video call, the called terminal sends a 200 OK response to the INVITE message.
[0047] Step 7: After the calling terminal receives the 200 OK (INVITE), it sends an ACK message to confirm the completion of session establishment and start bidirectional media transmission, i.e., officially start the holographic video call.
[0048] Step 8: Unified timestamp generation. The calling terminal generates the same timestamp for different video frames collected at the same time using the same clock source.
[0049] Step 9: Each video stream is transmitted using a separate RTP / RTCP transmission channel. During transmission, multiple RTP / RTCP transmission channels use space-time stamps (timestamp + camera position / identity) to synchronize video frames. The calling terminal adds timestamp and camera position / identity information to each video frame data. After receiving video frames from different RTP / RTCP transmission channels, the called terminal sorts and synchronizes the video frames according to the space-time stamp, ensuring that multiple perspective video streams remain consistent in time and present a coherent and synchronized picture effect. In space, the video streams are processed according to the video capture position or identity. Specifically: set the same timestamp field value in the RTP / RTCP packet header field, and write the camera position / identity in the Extension extension field in the RTP / RTCP packet header field.
[0050] In Figure 2 In the above-mentioned embodiments, the S-CSCF (Serving CSCF) represents a service call session control function, which is the central node of the IMS core network, responsible for processing the registration, authentication, session control (such as call establishment, hang-up) and service triggering of the UE (user terminal). The I-CSCF (Interrogating CSCF) represents an interrogating call session control function, which is usually used as the entry node of the operator network, responsible for querying the HSS (home subscriber server) to obtain the location information of the user, and routing the call to the correct S-CSCF. UE_A represents the calling terminal, and UE_B represents the called terminal.
[0051] In the above-mentioned embodiments, Figure 2 In the embodiments shown in the above-mentioned embodiments, no Internet holographic communication dedicated protocol, independent application or service needs to be introduced, and SIP protocol extension is directly based on the existing communication standard of the IMS network. The IMS network only needs to support the newly added SIP negotiation field (such as holographic capability reporting, media parameter negotiation, and Extension extension field of RTP / RTCP), without the need for large-scale modification of the infrastructure. The modification of the technical implementation is mainly concentrated on the terminal side equipment, which can significantly reduce the dependence on the IMS core network and the modification cost. Therefore, the scheme of the present application is compatible with the existing IMS network architecture, and can be quickly deployed in the operator's existing network environment.
[0052] In combination with the above embodiments, in an implementation manner, the step S103 can include: In the case that the response message does not carry the second holographic communication identifier indicating that the called terminal has holographic communication capability, or only carries the information of a single second transmission channel for receiving the multi-channel video stream of the calling terminal, a single-channel video call is established between the calling terminal and the called terminal.
[0053] In the present embodiment, after receiving the response message, the calling terminal specifically analyzes the content in the response message: if the response message does not contain the second holographic communication identifier indicating that the called terminal has holographic communication capability, or only contains the information of a single second transmission channel for receiving the multi-channel video stream of the calling terminal, it indicates that the type of the response message is the second response message. Then, the calling terminal directly establishes a single-channel video call with the called terminal.
[0054] In the present embodiment, when the called terminal receives the call request, if it confirms that it does not have holographic communication capability, it generates a second response message and carries the information of a single second transmission channel in the second response message. In this way, when the calling terminal determines that the type of the response message is the second response message, it will re-negotiate with the called terminal, and then establish a single-channel video call with the called terminal.
[0055] The following will describe how a calling terminal establishes a one-way video call with a called terminal in an embodiment of the present application. In this embodiment, the calling terminal has holographic communication capability, and the called terminal does not have holographic communication capability. Figure 3 is a flowchart of establishing a one-way video call between a calling terminal and a called terminal according to an embodiment of the present application. Referring to Figure 3 , the embodiment includes the following steps: Step 1: The calling terminal sends a call request, i.e. an INVITE message, to the called terminal. The INVITE message is sent using the SIP protocol, and the Contact: gfree-holography-video field (which is the first holographic communication identifier) is carried in the SIP header field, indicating that the calling terminal has holographic communication capability. In addition, information of a plurality of first transmission channels determined by the calling terminal (port, encoding format, resolution, etc.) is carried in the SDP message body.
[0056] Step 2: The called terminal returns a 183 Session Progress response to the calling terminal. The response message is sent using the SIP protocol. Since the called terminal does not have holographic communication capability, the holographic communication capability identifier is removed from the SIP header field. In addition, information of only one second transmission channel determined by the called terminal (port, encoding format, resolution, etc.) is carried in the SDP message body, i.e. the original encoding mode, resolution configuration, etc., and the a line corresponding to the transmission channel carries a=content: visualized-voice-call, indicating that only 2D video call is supported.
[0057] An example of the SIP header field of the 183 Session Progress is as follows: "SIP / 2.0 183 Session Progress Allow: INVITE, ACK, OPTIONS, CANCEL, BYE, UPDATE, INFO, REFER, NOTIFY, MESSAGE, PRACK Contact: +g.3gpp.icsi-ref="urn%3Aurn-7%3A3gpp-service.ims.icsi.mmtel"; +sip.instance=" <urn:gsma:imei:35293609-715199-0> 183 An example of an SDP message body for Session Progress is as follows: # Ordinary 2D video call (H.264 encoding, 720p 30 frames): # Declare this is a video stream, using port 64084 and the RTP / AVP protocol, with a payload type of 118 m=video 64084 RTP / AVP 118 # Specify that payload type 118 corresponds to H.264 encoding, with a clock rate of 90000 Hz a=rtpmap:118 H264 / 90000 # Define additional parameters for payload type 118: maximum bitrate of 2174 kbps, profile and level for H.264 (42 supports low delay, low complexity; C0 resolution dependent), packetization mode 1 (suitable for real-time communication), vps and sps and pps for video frame-related configuration a=fmtp:118 max-br=2174;profile-level-id=42C01F;packetization-mode=1;sprop-parameter-sets= <video parameter sets data> # frame rate a=framerate:30 # identify that the video stream belongs to the visualized voice call service a=content:visualized-voice-call” Step 3: After the calling terminal receives the 183 Session Progress, the SIP PRACK (Provisional Acknowledgement) is sent to confirm the media parameters.
[0058] Step 4: After the called terminal receives the PRACK message, a 200 OK response PRACK message is returned.
[0059] Step 5: The calling terminal sends an UPDATE message to negotiate a 2D video call according to the media capabilities supported by the called terminal, removes the holographic communication capability identifier in the SIP header field, removes the holographic media negotiation related flag in the SDP message body, and only carries a=content: visualized-voice-call to indicate that only a 2D video call is negotiated.
[0060] Step 6: The called terminal sends a 200 OK response UPDATE message.
[0061] Step 7: The called terminal starts ringing and sends a 180 Ringing response message.
[0062] Step 8: After the called terminal answers the call, a 200 OK response INVITE message is sent to the calling terminal.
[0063] Step 9: When the calling terminal receives the final response message 200 OK (INVITE) from the called terminal, an ACK message is sent to confirm that the session establishment is completed and the bidirectional media transmission is started, that is, the 2D video call is formally started.
[0064] The embodiment provides an emergency measure when the called terminal does not have holographic communication capability, so that the called terminal can perform a video call with the calling terminal as long as it has a normal camera, and the influence on the user's video call experience can be avoided.
[0065] In combination with the above embodiment, in an implementation manner, after step S103, the method of the application can further include: obtaining a network state change of each second transmission channel; if the amplitude of the network state change is greater than a preset amplitude, updating the information of the plurality of first transmission channels and the information of the plurality of second transmission channels.
[0066] In the embodiment, the network jitter monitoring module is arranged in the calling terminal. The network jitter monitoring module monitors the RTP / RTCP feedback information received by each second transmission channel in real time, and determines the information of each index of each second transmission channel according to the RTP / RTCP feedback information. Then, the information of each transmission channel is re-negotiated and updated according to the change range of part of the indexes.
[0067] In an embodiment, for a target second transmission channel, if the change range of one or more indexes of all indexes corresponding to the target second transmission channel is greater than a pre-set corresponding range, it is determined that the information of the plurality of first transmission channels and the information of the plurality of second transmission channels need to be updated. The target second transmission channel is part of the second transmission channels, and the target second transmission channel can be set according to actual needs.
[0068] In an embodiment, the network state can include indexes such as network bandwidth, round-trip time, packet loss rate, and delay. The calling terminal can accurately perceive the network state by monitoring the network bandwidth, round-trip time, packet loss rate, and delay in real time through the network jitter monitoring module. When the network state is not good, the update of the information of each transmission channel is automatically triggered, for example, the resolution and frame rate of each transmission channel can be reduced, or a more efficient encoding mode can be switched to reduce the amount of data transmission, thereby prioritizing the smooth progress of the video call. When the network state improves, the update of the information of each transmission channel is triggered again, for example, high resolution and high frame rate can be restored to improve video quality, thereby improving the quality of the video call. Such adaptive adjustment capability enables the video call to maintain good usability and quality in complex and variable network environments and different device performance conditions, and can significantly enhance the environmental adaptability and stability of the video call.
[0069] In combination with the above embodiments, in an embodiment, updating the information of the plurality of first transmission channels and the information of the plurality of second transmission channels includes: updating the information of the plurality of first transmission channels according to the change of the network state of each second transmission channel; sending a parameter update request to the called terminal, the parameter update request carrying the updated information of the plurality of first transmission channels, and the parameter update request being used to instruct the called terminal to update the information of the plurality of second transmission channels according to the updated information of the plurality of first transmission channels.
[0070] In this embodiment, the calling terminal can first update the information of the plurality of first transmission channels, and then send a parameter update request to the called terminal according to the updated information of the plurality of first transmission channels. After receiving the parameter update request, the called terminal updates the information of the plurality of second transmission channels according to the updated information of the plurality of first transmission channels in the parameter update request, and obtains the updated information of the plurality of second transmission channels. Then, the called terminal returns a response to the parameter update request to the calling terminal, and the response carries the updated information of the plurality of second transmission channels. After that, the calling terminal can receive the video stream transmitted by the called terminal based on the updated information of the plurality of first transmission channels, and the called terminal can receive the video stream transmitted by the calling terminal based on the updated information of the plurality of second transmission channels.
[0071] In this embodiment, when updating the information of each transmission channel, the calling terminal first updates the information of the plurality of first transmission channels, and then the called terminal updates the information of the plurality of second transmission channels according to the updated information of the plurality of first transmission channels. In this way, the updating efficiency of the information of the transmission channel can be effectively improved.
[0072] In the following, an embodiment will be described to illustrate how to update the information of the transmission channel in the holographic video call. In this embodiment, the monitored indicators are the packet loss rate and the delay. Figure 4 is a flowchart of updating the information of the transmission channel according to an embodiment of the present application. Referring to Figure 4 , the embodiment includes the following steps: Step 1: The calling terminal receives the RTP / RTCP feedback information of each second transmission channel through the network jitter monitoring module, and obtains the packet loss rate and the delay of each transmission channel according to the feedback information. Then, the calling terminal re-negotiates the information of the transmission channel according to the change range (increase range or decrease range) of the packet loss rate and the change range (increase range or decrease range) of the delay.
[0073] Step 2: If it is necessary to adjust (increase or decrease) the resolution, the calling terminal re-negotiates the resolution of the transmission channel through the SIP re-INVITE message. The header field information of the SIP and the content of the SDP message body are basically the same as those in the INVITE message in Figure 2 , and only the corresponding resolution in the SDP message body is different.
[0074] Step 3: The called terminal returns a 200 OK response. In the response, the header field information of the SIP and the content of the SDP message body are the same as those in the 183 Session Progress in Figure 2 , and only the resolution is different.
[0075] Step 4: The calling terminal sends an ACK message to confirm that the session establishment is completed, and starts the media transmission.
[0076] The above detailed the updating scheme of the information of each transmission channel in the holographic video call scenario. The updating scheme is also applicable to the transmission channel in the single-channel video call scenario, and the updating principles are the same. Details are not described herein again.
[0077] The method for establishing a video call of the present application will be described from the side of the called terminal. Figure 5 FIG. 2 is a flowchart of another method for establishing a video call according to an embodiment of the present application. Referring to FIG. 2, Figure 5 The method for establishing a video call of the present application can include: In step S201, a call request sent by a calling terminal is received through an IMS network. The call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a multi-channel video stream of the called terminal. In step S202, a response message is generated based on the call request. In step S203, the response message is sent to the calling terminal through the IMS network. The response message is used to instruct the calling terminal to establish a video call with the called terminal.
[0078] According to the method of the present application, the called terminal first receives a call request sent by a calling terminal through an IMS network. The call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a multi-channel video stream of the called terminal. Then, the called terminal generates a response message based on the call request and sends the response message to the calling terminal, so that the calling terminal establishes a video call with the called terminal based on the type of the response message. According to the method of the present application, the calling terminal first initiates a call request aiming to establish a holographic video call, and the called terminal generates a response message based on its own capability information. The type of the response message determines the type of the video call to be finally established. In this way, the terminal can not only perform holographic video call under the IMS network, but also perform single-channel video call under the IMS network. The video call capability of the terminal under the IMS network is greatly improved, and the problem that the terminal under the IMS network can only use a single camera to collect a single-channel video stream in the related art, which cannot meet the user's demand for immersive multi-view holographic video call, is overcome.
[0079] In combination with the above embodiments, in an implementation, step S202 includes: In a case where it is determined based on the call request that the called terminal has holographic call capability, information of a plurality of second transmission channels for receiving a multi-channel video stream of the calling terminal is determined according to the information of the plurality of first transmission channels. generate a first response message, wherein the first response message carries second holographic communication identification indicating that the called terminal has holographic communication capability and information of a plurality of second transmission channels.
[0080] In the embodiment, after receiving the call request, if it is determined that the call request carries first holographic communication identification indicating that the calling terminal has holographic communication capability and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal, the called terminal continues to determine whether the called terminal has holographic communication capability. If it is determined that the called terminal has holographic communication capability, information of a plurality of second transmission channels for receiving a plurality of video streams of the calling terminal is determined according to the information of the plurality of first transmission channels, and then a first response message is generated according to second holographic communication identification indicating that the called terminal has holographic communication capability and information of the plurality of second transmission channels.
[0081] In the embodiment, the called terminal generates the first response message when it is determined that the called terminal has holographic communication capability, which can ensure successful establishment of holographic video call between the calling terminal and the called terminal.
[0082] In combination with the above embodiments, in an implementation, step S202 can include: determining information of a single second video transmission channel for receiving a video stream of the calling terminal, in a case that it is determined that the called terminal does not have holographic call capability based on the call request; generating a second response message, wherein the second response message carries information of the single second video transmission channel.
[0083] In the embodiment, after receiving the call request, if it is determined that the call request carries first holographic communication identification indicating that the calling terminal has holographic communication capability and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal, the called terminal continues to determine whether the called terminal has holographic call capability. If it is determined that the called terminal does not have holographic communication capability, a second response message is directly generated according to information of a transmission channel supported by the called terminal.
[0084] In the embodiment, the called terminal generates the second response message when it is determined that the called terminal does not have holographic communication capability, which can ensure successful establishment of single-channel video call between the calling terminal and the called terminal.
[0085] In combination with the above embodiments, in an implementation, after step S203, the method of the present application can further include: in a case that a message indicating that the holographic call is successfully established and sent by the calling terminal is received, a plurality of video streams collected by a camera of the called terminal, and information of each video frame corresponding to a timestamp of each video stream and the camera are acquired; According to the information of the plurality of first transmission channels, the multi-path video stream, the timestamps corresponding to each video frame, and the information of the cameras are transmitted to the calling terminal.
[0086] In this embodiment, after the called terminal establishes the holographic video call with the calling terminal, the called terminal can transmit the multi-path video stream to the calling terminal. The called terminal also has a unified timestamp generation module, which can mark the same timestamp for different video frames collected at the same time when collecting multi-angle video streams through multiple cameras of the called terminal. Then, the called terminal transmits the multi-path video stream marked with the timestamp and the information of the cameras of each video frame to the calling terminal according to the information of the plurality of first transmission channels. In this way, the calling terminal can use the received unified timestamp to accurately time-align the multi-path video stream collected at the same time, ensure that the pictures of all angles are strictly synchronized, and prevent delays or asynchronization between different angles. At the same time, through the information of the cameras, the calling terminal can accurately identify the spatial source corresponding to each video stream, so as to sort, reconstruct and render these pictures according to the correct spatial relationship, and finally realize a coherent, immersive and spatially accurate holographic multi-view call experience.
[0087] The information of the cameras includes the identity of the cameras, the positions of the cameras, etc., which can be set according to actual needs.
[0088] In this embodiment, according to the information of the plurality of first transmission channels, the multi-path video stream, the timestamps corresponding to each video frame, and the information of the cameras are transmitted to the calling terminal, which can effectively enhance the holographic video call experience between the calling terminal and the called terminal.
[0089] In combination with the above embodiments, in one implementation, in a case where a message indicating that the holographic call is successfully established is received from the calling terminal, the method of the present application further includes: Step S1: In the process of receiving the multi-path video stream sent by the calling terminal, if there is a missing video frame, determining the target timestamp corresponding to the missing video frame and the target video stream to which the missing video frame belongs.
[0090] In this embodiment, an edge cloud video prediction gateway is arranged on the edge side of the called terminal, and the called terminal can specifically execute steps S1-S5 through the edge cloud video prediction gateway. If there is a case where a video stream is missing a video frame in a certain angle, the edge cloud video prediction gateway can complete the image information of the missing frame through AI machine learning.
[0091] For example, the called terminal receives the 3-way video stream sent by the calling terminal, which is the video stream under view 1, the video stream under view 2 and the video stream under view 3 respectively. When the called terminal detects that a certain video frame is missing, first determine the target timestamp corresponding to the missing video frame (i.e. the timestamp of the other video frame time-synchronized therewith) and the video stream to which the missing video frame belongs, for example, the video frame under view 3.
[0092] Step S2: Obtain the first video frame and the second video frame related to the target timestamp in each of the video streams other than the target video stream, the first video frame being the video frame corresponding to the target timestamp, and the second video frame being the previous video frame of the first video frame.
[0093] In the above example, the first video frame Pv11 and the second video frame Pv12 related to the target timestamp under view 1, and the first video frame Pv21 and the second video frame Pv22 related to the target timestamp under view 2 are obtained. Pv12 is the previous video frame received before Pv11, and Pv22 is the previous video frame received before Pv21.
[0094] Step S3: input the first video frame and the second video frame into the feature pyramid network to obtain a plurality of feature maps of different resolutions corresponding to each of the video streams other than the target video stream.
[0095] In the above example, Pv11, Pv12, Pv21 and Pv22 are input into the feature pyramid network, and the feature pyramid network finally obtains a plurality of feature maps of different resolutions through different degrees of downsampling and convolution operations on the images. For example, Pv11, Pv12, Pv21 and Pv22 correspond to two kinds of feature maps of a first resolution (high) and a second resolution (low) respectively, the first resolution being greater than the second resolution.
[0096] Step S4: perform optical flow calculation on the first video frame and the second video frame to obtain displacement information of the pixel points of the video frames corresponding to each of the video streams other than the target video stream.
[0097] In the above example, optical flow calculation is performed on Pv11 and Pv12 to obtain the displacement information S1 of the pixel points of the video frames corresponding to the video stream under view 1; and optical flow calculation is performed on Pv21 and Pv22 to obtain the displacement information S2 of the pixel points of the video frames corresponding to the video stream under view 2.
[0098] Step S5: determine the image information of the missing video frame according to the plurality of feature maps of different resolutions and the displacement information.
[0099] According to the first resolution and the second resolution of the feature maps corresponding to Pv11, Pv12, Pv21, and Pv22 respectively, and the displacement information S1 and the displacement information S2, the image information of the missing video frame is determined.
[0100] In the present embodiment, the feature pyramid network constructs a multi-resolution feature map set by performing different degrees of downsampling and convolution operations on the video frames. It captures global structural features at low resolution and obtains detailed texture features at high resolution. The optical flow calculation uses the brightness consistency and small motion assumption of pixels between adjacent video frames to describe the motion information of the pixel points in the image by calculating the displacement vectors of the pixels. The video prediction gateway inputs the extracted feature maps and motion information of different resolutions into the prediction model based on the deep learning algorithm, uses the existing picture features and motion trends of other perspectives to infer and reconstruct the picture content of the missing video frame, fills in the missing video frame, and thus ensures the integrity and continuity of the video picture, effectively solving the picture missing problem caused by network jitter.
[0101] In combination with the above embodiments, in an implementation manner, the step S5 can include: The step S51 includes determining the target displacement information corresponding to the missing video frame according to each displacement information.
[0102] In combination with the above examples in the steps S1-S5, the target displacement information S is determined according to the displacement information S1 and S2. The displacement information is represented by a matrix, and when the target displacement information is determined, the mean value of the corresponding elements in S1 and S2 can be calculated to obtain a new matrix, which is the target displacement information S. Alternatively, the values of the corresponding elements in S1 and S2 can be weighted and summed to obtain the target displacement information S, wherein the weight corresponding to S1 represents the influence degree (the higher the similarity degree, the greater the influence degree) of the video frame under the perspective 1 corresponding to the video frame under the perspective 3. The weight corresponding to S2 represents the influence degree (the higher the similarity degree, the greater the influence degree) of the video frame under the perspective 2 corresponding to the video frame under the perspective 3. The weight corresponding to S1 and the weight corresponding to S2 can be calibrated in advance.
[0103] The step S52 includes determining the candidate video frame corresponding to the missing video frame for each video stream except the target video stream according to the corresponding feature maps of different resolutions, the displacement information, and the target displacement information.
[0104] With the foregoing example, first, the multi-resolution feature map corresponding to the view 1 generated in step S3 (containing high-resolution texture features and low-resolution structure features) and the displacement information S1 of the view 1 calculated in step S4 are input into the pre-trained image generation network together with the target displacement information S determined in step S51. The network uses the target displacement information S to perform spatial transformation and feature resampling (Feature Warping) on the feature map of the view 1, warps the features of the view 1 to the predicted position of the view 3, and thus generates a first candidate video frame Cv1. Cv1 represents the current frame image of the view 3 inferred only according to the picture information and motion trend of the view 1. Then, the multi-resolution feature map corresponding to the view 2, the displacement information S2 of the view 2, and the target displacement information S are input into the image generation network. The image generation network also performs corresponding spatial mapping and texture reconstruction on the feature map of the view 2 according to the target displacement information S, and generates a second candidate video frame Cv2. Cv2 represents the current frame image of the view 3 inferred only according to the picture information and motion trend of the view 2.
[0105] In this embodiment, inputting the displacement information S1 and S2 into the image generation network can provide the model with the pixel motion prior and spatio-temporal evolution rule within the reference view. When the image generation network processes the target displacement information S and the multi-resolution feature map, the source displacement information can assist the model in accurately identifying the occlusion relationship and object boundary in the dynamic scene, so as to achieve more accurate motion compensation in the feature resampling and spatial mapping process. This mechanism can effectively avoid the pixel misplacement or artifacts that may be caused by simply relying on the target displacement S, ensure that the generated candidate video frame is more consistent with the real physical motion logic in terms of dynamic texture and structure layout, and significantly improve the clarity and continuity of the missing frame reconstruction.
[0106] Step S53, determining the image information of the missing video frame according to the candidate video frames corresponding to the video streams other than the target video stream.
[0107] In this embodiment, the image information of the candidate video frame is represented by a matrix. With the foregoing example, when determining the image information of the missing video frame, the mean of the corresponding elements in Cv1 and Cv2 can be calculated to obtain a new matrix, which is the image information Cv of the missing video frame. Alternatively, the values of the corresponding elements in Cv1 and Cv2 can be weighted and summed to obtain the image information Cv of the missing video frame, wherein the weight corresponding to Cv1 is positively correlated with the spatial position correlation (or the confidence of network prediction) between the view 1 and the view 3, and the weight corresponding to Cv2 is positively correlated with the spatial position correlation (or the confidence of network prediction) between the view 2 and the view 3.
[0108] An embodiment will be described below on how to update the information of transmission channels in holographic video call. In this embodiment, the monitored indicators are packet loss rate and delay. Figure 6 is a schematic diagram of a missing video frame completion process according to an embodiment of the present application. Referring to Figure 6 , the embodiment includes the following steps: The edge cloud video prediction gateway monitors the video streams received by the called terminal in real time through the video stream monitoring module; when it is detected that there are missing video frames due to network jitter, video frames are extracted from other perspectives, and the image information of the missing video frames is completed through the feature pyramid network and optical flow calculation; finally, the completed complete video streams are transmitted to the called terminal.
[0109] In this embodiment, by introducing the edge cloud video prediction gateway, high-efficiency fault tolerance can be achieved in combination with AI deep learning algorithm. When the edge cloud video prediction gateway detects that the video frames of a certain perspective are missing, first, the feature pyramid network is used to extract multi-level features from the video frames under other normal perspectives, to capture image information of different resolutions; then, by means of optical flow prediction technology, the motion trend and object change law between video frames are analyzed, so as to accurately predict the picture content of the missing video frames. In this way, the image information of the missing video frames can be quickly and effectively recovered, the integrity and continuity of the video picture are ensured, and the negative impact of picture loss on the call effect is greatly reduced.
[0110] The establishment device of the video call provided by the embodiment of the present application will be described below. The establishment device of the video call described below can be correspondingly referred to the establishment method of the video call described above.
[0111] The present application first provides an establishment device 700 of a video call, applied to a calling terminal. Figure 7 is a structural block diagram of an establishment device of a video call according to an embodiment of the present application. Referring to Figure 7 , the establishment device 700 can include: A sending module 701 is configured to send a call request to a called terminal through an IMS network, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal. A receiving module 702 is configured to receive a response message sent by the called terminal through the IMS network, wherein the response message is generated by the called terminal based on the call request. An establishment module 703 is configured to establish a video call with the called terminal based on the response message.
[0112] According to the application, the establishing device 700 for video call is provided, and the establishing module 703 is specifically configured to: in the case that the second holographic communication identifier indicating that the called terminal has holographic communication capability is carried in the response message, and the information of the multiple second transmission channels for receiving the multiple video streams of the calling terminal is carried, establish a holographic video call with the called terminal; and the information of the multiple second transmission channels is determined based on the information of the multiple first transmission channels.
[0113] According to the application, the establishing device 700 for video call further comprises an obtaining module configured to: obtain the multiple video streams collected by the camera of the calling terminal, and the information of each video frame corresponding to a timestamp and the camera; and transmit the multiple video streams, the information of each video frame corresponding to a timestamp and the camera to the called terminal according to the information of the multiple second transmission channels.
[0114] According to the application, the establishing device 700 for video call is provided, and the establishing module 703 is specifically configured to: in the case that the second holographic communication identifier indicating that the called terminal has holographic communication capability is carried in the response message, and the information of the multiple second transmission channels for receiving the multiple video streams of the calling terminal is carried, establish a holographic video call with the called terminal; and the information of the multiple second transmission channels is determined based on the information of the multiple first transmission channels.
[0115] According to the application, the establishing device 700 for video call further comprises an updating module configured to: obtain the network state change of each second transmission channel; and update the information of the multiple first transmission channels and the information of the multiple second transmission channels if the amplitude of the network state change is greater than a preset amplitude.
[0116] According to the application, the updating module of the establishing device 700 for video call is specifically configured to: update the information of the multiple first transmission channels according to the network state change of each second transmission channel; and send a parameter update request to the called terminal, wherein the updated information of the multiple first transmission channels is carried in the parameter update request, and the parameter update request is used to instruct the called terminal to update the information of the multiple second transmission channels according to the updated information of the multiple first transmission channels.
[0117] According to the application, the sending module 701 of the establishing device 700 for video call is specifically configured to: send a call request to the called terminal through the SIP protocol, wherein the first holographic communication identifier is located in the SIP header field of the call request, and the information of the multiple first transmission channels is located in the SDP message body of the call request.
[0118] The application first provides another video call establishment device 800 applied to a called terminal. Figure 8 is a structural block diagram of another video call establishment device according to an embodiment of the application. Referring to Figure 7 , the establishment device 800 can include: a receiving module 801 configured to receive a call request sent by a calling terminal through an IMS network, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; a generating module 802 configured to generate a response message based on the call request; a sending module 803 configured to send the response message to the calling terminal through the IMS network, wherein the response message is used to instruct the calling terminal to establish a video call with the called terminal.
[0119] According to the video call establishment device 800 provided by the application, the generating module 802 is specifically configured to: in a case where it is determined based on the call request that the called terminal has holographic call capability, determine information of a plurality of second transmission channels for receiving a plurality of video streams of the calling terminal according to the information of the plurality of first transmission channels; and generate a first response message, wherein the first response message carries a second holographic communication identifier indicating that the called terminal has holographic communication capability and the information of the plurality of second transmission channels.
[0120] According to the video call establishment device 800 provided by the application, the generating module 802 is specifically configured to: in a case where it is determined based on the call request that the called terminal does not have holographic call capability, determine information of a single second video transmission channel for receiving a video stream of the calling terminal; and generate a second response message, wherein the second response message carries the information of the single second video transmission channel.
[0121] According to the video call establishment device 800 provided by the application, the device further includes a transmission module configured to: in a case where a message indicating that holographic call establishment is successful sent by the calling terminal is received, acquire a plurality of video streams collected by a camera of the called terminal and information of each video frame corresponding to a timestamp of each video stream and the camera; and transmit the plurality of video streams, the information of each video frame corresponding to the timestamp of each video stream and the camera to the calling terminal according to the information of the plurality of first transmission channels.
[0122] The video call establishment apparatus 800 provided in this application further includes: a completion module, configured to: in the process of receiving multiple video streams sent by the calling terminal, if there are missing video frames, determine the target timestamp and the target video stream to which the missing video frame belongs; obtain a first video frame and a second video frame related to the target timestamp from each of the other video streams besides the target video stream, wherein the first video frame is the video frame corresponding to the target timestamp and the second video frame is the previous video frame of the first video frame; input the first video frame and the second video frame into a feature pyramid network to obtain multiple feature maps of different resolutions corresponding to each of the other video streams besides the target video stream; perform optical flow calculation on the first video frame and the second video frame to obtain the pixel displacement information of the video frames corresponding to each of the other video streams besides the target video stream; and determine the image information of the missing video frame based on the multiple feature maps of different resolutions and the displacement information.
[0123] According to the video call establishment apparatus 800 provided in this application, the completion module is specifically used for: determining the target displacement information corresponding to the missing video frame based on each displacement information; determining the candidate video frame corresponding to the missing video frame for each video stream other than the target video stream based on the corresponding feature map of different resolutions, displacement information and the target displacement information; and determining the image information of the missing video frame based on the candidate video frames corresponding to each video stream other than the target video stream.
[0124] Figure 9 This is a schematic diagram of the physical structure of an electronic device shown in an embodiment of this application, such as... Figure 9 As shown, the electronic device may include a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 can call a computer program in the memory 930 to execute the steps of a video call establishment method.
[0125] Further, the logic instructions in the memory 930 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0126] In another aspect, the embodiments of the present application also provide a computer program product, which comprises a computer program, the computer program can be stored in a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the steps of the video call establishing method provided by the above-mentioned embodiments.
[0127] In another aspect, the embodiments of the present application also provide a processor readable storage medium, which stores a computer program, and the computer program is used to make the processor execute the steps of the video call establishing method provided by the above-mentioned embodiments.
[0128] The processor readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to a magnetic storage (such as a floppy disk, a hard disk, a magnetic tape, a magneto-optical disk (MO), etc.), an optical storage (such as a CD, a DVD, a BD, a HVD, etc.), and a semiconductor memory (such as a ROM, an EPROM, an EEPROM, a NAND FLASH, a solid state disk (SSD)), etc.
[0129] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement it without creative labor.
[0130] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and the necessary general hardware platform from the above description of the embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that contributes to the technical solutions can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the methods.
[0131] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.< / video> < / video> < / video> < / video> < / video> < / video> < / video>
Claims
1. A method of establishing a video call, the method comprising: Applied to a calling terminal, the method comprises: sending, through an IMS network, a call request to a called terminal, the call request carrying a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; receiving, through the IMS network, a response message sent by the called terminal, the response message being generated by the called terminal based on the call request; based on the response message, establishing a video call with the called terminal.
2. The method of claim 1, wherein, The establishing a video call with the called terminal based on the response message comprises: in a case where the response message carries a second holographic communication identifier indicating that the called terminal has holographic communication capability, and information of a plurality of second transmission channels for receiving a plurality of video streams of the calling terminal, establishing a holographic video call with the called terminal; wherein the information of the plurality of second transmission channels is determined based on the information of the plurality of first transmission channels.
3. The method of claim 2, wherein, After the establishing a video call with the called terminal, the method further comprises: obtaining a plurality of video streams collected by a camera of the calling terminal, and information of each video frame corresponding to a timestamp and the camera; transmitting the plurality of video streams, the information of each video frame corresponding to a timestamp and the camera, to the called terminal according to the information of the plurality of second transmission channels.
4. The method of claim 1, wherein, The establishing a video call with the called terminal based on the response message comprises: in a case where the response message does not carry a second holographic communication identifier indicating that the called terminal has holographic communication capability, or only carries information of a single second transmission channel for receiving a video stream of the calling terminal, establishing a single-channel video call with the called terminal.
5. The method of claim 2, wherein, After the establishing a video call with the called terminal, the method further comprises: obtaining a network state change of each second transmission channel; if a magnitude of the network state change is greater than a preset magnitude, updating the information of the plurality of first transmission channels and the information of the plurality of second transmission channels.
6. The method of claim 5, wherein, The updating the information of the plurality of first transmission channels and the information of the plurality of second transmission channels comprises: updating the information of the plurality of first transmission channels according to the network state change of each second transmission channel; sending a parameter update request to the called terminal, the parameter update request carrying the updated information of the plurality of first transmission channels, the parameter update request being used to instruct the called terminal to update the information of the plurality of second transmission channels according to the updated information of the plurality of first transmission channels.
7. The method according to any one of claims 1 to 6, characterized in that, The sending a call request to the called terminal comprises: sending a call request to the called terminal through a SIP protocol, wherein the first holographic communication identifier is located in a SIP header field of the call request, and the information of the plurality of first transmission channels is located in an SDP message body of the call request.
8. A method of establishing a video call, the method comprising: Applied to a called terminal, the method comprises: Receiving, by an IMS network, a call request sent by a calling terminal, the call request carrying a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels for receiving a plurality of video streams of the called terminal; Generating a response message based on the call request; Sending, by the IMS network, the response message to the calling terminal, the response message being used to instruct the calling terminal to establish a video call with the called terminal.
9. The method of claim 8, wherein, The generating of the response message based on the call request comprises: In a case where it is determined based on the call request that the called terminal has holographic call capability, determining information of a plurality of second transmission channels for receiving a plurality of video streams of the calling terminal according to the information of the plurality of first transmission channels; Generating a first response message, the first response message carrying a second holographic communication identifier indicating that the called terminal has holographic communication capability and the information of the plurality of second transmission channels.
10. The method of claim 8, wherein, The generating of the response message based on the call request comprises: In a case where it is determined based on the call request that the called terminal does not have holographic call capability, determining information of a single second video transmission channel for receiving a video stream of the calling terminal; Generating a second response message, the second response message carrying the information of the single second video transmission channel.
11. The method of claim 8, wherein, After the sending of the response message to the calling terminal, the method further comprises: In a case where a message indicating that the holographic call is successfully established is received from the calling terminal, obtaining a plurality of video streams captured by a camera of the called terminal, and information of each video frame corresponding to a timestamp and the camera in each video stream; Transmitting, according to the information of the plurality of first transmission channels, the plurality of video streams, the information of each video frame corresponding to a timestamp and the camera, to the calling terminal.
12. The method according to any one of claims 8-11, characterized in that, In a case where a message indicating that the holographic call is successfully established is received from the calling terminal, the method further comprises: In a process of receiving the plurality of video streams sent by the calling terminal, if there is a missing video frame, determining a target timestamp corresponding to the missing video frame and a target video stream to which the missing video frame belongs; Obtaining a first video frame and a second video frame related to the target timestamp in each video stream except the target video stream, the first video frame being a video frame corresponding to the target timestamp, and the second video frame being a previous video frame of the first video frame; Inputting the first video frame and the second video frame into a feature pyramid network to obtain a plurality of feature maps of different resolutions corresponding to each video stream except the target video stream; Performing optical flow calculation on the first video frame and the second video frame to obtain displacement information of pixel points of the video frame corresponding to each video stream except the target video stream; Determining image information of the missing video frame according to the plurality of feature maps of different resolutions and the displacement information.
13. The method of claim 12, wherein, The determining of the image information of the missing video frame according to the plurality of feature maps of different resolutions and the displacement information comprises: According to each displacement information, determine the target displacement information corresponding to the missing video frame; For each video stream except the target video stream, according to the corresponding feature map with different resolutions, displacement information and the target displacement information, determine the candidate video frame corresponding to the missing video frame; According to the candidate video frame corresponding to each video stream except the target video stream, determine the image information of the missing video frame.
14. An apparatus for establishing a video call, the apparatus comprising: The device is deployed in a calling terminal, and the device comprises: The sending module is configured to send a call request to a called terminal through an IMS network, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels used for receiving a plurality of video streams of the called terminal; The receiving module is configured to receive a response message sent by the called terminal through the IMS network, wherein the response message is generated by the called terminal based on the call request; The establishing module is configured to establish a video call with the called terminal based on the response message.
15. An apparatus for establishing a video call, the apparatus comprising: The device is deployed in a calling terminal, and the device comprises: The receiving module is configured to receive a call request sent by a calling terminal through an IMS network, wherein the call request carries a first holographic communication identifier indicating that the calling terminal has holographic communication capability, and information of a plurality of first transmission channels used for receiving a plurality of video streams of the called terminal; The generating module is configured to generate a response message based on the call request; The sending module is configured to send the response message to the calling terminal through the IMS network, wherein the response message is used to instruct the calling terminal to establish a video call with the called terminal.
16. An electronic device comprising a processor and a memory having a computer program stored therein, characterized in that The processor executes the computer program to implement the steps of the video call establishment method of any one of claims 1 to 7, or implement the steps of the video call establishment method of any one of claims 7 to 13.
17. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the video call establishment method of any one of claims 1 to 7, or implement the video call establishment method of any one of claims 8 to 13.
18. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the video call establishment method of any one of claims 1 to 7, or implement the steps of the video call establishment method of any one of claims 8 to 13.