Communication method, ims network element, storage medium and computer program product

CN122601649APending Publication Date: 2026-08-18ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610964288.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本申请的主要目的在于提供一种通信方法、IMS网元、存储介质及计算机程序产品,旨在解决如何在IMS中实现生成式通信的技术问题

Benefits of technology

[0010] This application provides a communication method, IMS network element, storage medium, and computer program product, relating to the field of communication technology. The method includes: receiving a first SIP invitation request from a first terminal, the first SIP invitation request being a request from the first terminal to initiate a video call to a second terminal; responding to the first SIP invitation request and determining that the first terminal has enabled semantic service functions and possesses video semantic extraction capabilities, sending a first SIP invitation response to the first terminal, the first SIP invitation response carrying a first SDP response, the first SDP response including first information, the first information being used to instruct the first terminal to convert video data collected during the video call into video semantic features; receiving the video semantic features sent by the first terminal, and audio data collected by the first terminal during the video call; performing timestamp alignment based on the video semantic features and audio data, and generating first video call data based on the timestamp-aligned video semantic features and audio data, and sending it to the second terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601649A_ABST
    Figure CN122601649A_ABST
Patent Text Reader

Abstract

The application discloses a communication method, an IMS network element, a storage medium and a computer program product, relates to the technical field of communication, and the method comprises the following steps: receiving a first SIP invitation request from a first terminal, the first SIP invitation request being a request of initiating a video call by the first terminal to a second terminal; in response to the first SIP invitation request, determining that the first terminal opens a semantic service function and has video semantic extraction capability, sending a first SIP invitation response carrying a first SDP response to the first terminal, the first SDP response comprising first information used for instructing the first terminal to convert video data collected in the video call into video semantic features; performing timestamp alignment based on video semantic features and audio data sent by the first terminal, and generating first video call data based on the video semantic features and the audio data after the timestamp alignment, and sending the first video call data to the second terminal. The application can realize generative communication in the IMS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to communication methods, IMS network elements, storage media and computer program products. Background Technology

[0002] With the evolution of multimedia communication technologies, the Internet Protocol Multimedia Subsystem (IMS), as the core architecture providing IP (Internet Protocol) multimedia services, is widely used in services such as video calls and instant messaging. In existing IMS video call scenarios, communication mainly relies on the real-time acquisition and transmission of real audio and video streams. That is, the sending end captures local camera footage and microphone audio, encodes it, and transmits it over the network to the receiving end for decoding and playback. This method can reproduce a realistic communication scenario and is currently the mainstream implementation form of video communication.

[0003] However, with the constraints of bandwidth resources and the increasing demands of users for privacy protection and personalized expression, traditional real-time video streaming methods are gradually revealing their limitations. On the one hand, high-definition video streams require high network bandwidth and are prone to stuttering or image quality degradation in weak network environments; on the other hand, some users do not wish to expose their real physical environment or facial appearance during video calls.

[0004] Therefore, how to implement generative communication in IMS has become an important research direction in the field of multimedia communication. Summary of the Invention

[0005] The main purpose of this application is to provide a communication method, IMS network element, storage medium and computer program product, which aims to solve the technical problem of how to implement generative communication in IMS.

[0006] To achieve the above objectives, this application provides a communication method, comprising: Receive a first SIP (Session Initiation Protocol) invitation request from the first terminal, the first SIP invitation request being a request from the first terminal to initiate a video call to the second terminal; In response to the first SIP invitation request, which determines that the first terminal has enabled semantic service functions and has video semantic extraction capabilities, a first SIP invitation response is sent to the first terminal. The first SIP invitation response carries a first SDP (Session Description Protocol) response. The first SDP response includes first information, which is used to instruct the first terminal to convert the video data it collects in the video call into video semantic features. Receive the video semantic features sent by the first terminal, and the audio data collected by the first terminal during the video call; Based on the video semantic features and the audio data, timestamp alignment is performed, and based on the timestamp-aligned video semantic features and audio data, first video call data is generated and sent to the second terminal.

[0007] In addition, to achieve the above objectives, this application also provides an IMS network element, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the communication method described above.

[0008] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the communication method described above.

[0009] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the communication method described above.

[0010] This application provides a communication method, IMS network element, storage medium, and computer program product, relating to the field of communication technology. The method includes: receiving a first SIP invitation request from a first terminal, the first SIP invitation request being a request from the first terminal to initiate a video call to a second terminal; responding to the first SIP invitation request and determining that the first terminal has enabled semantic service functions and possesses video semantic extraction capabilities, sending a first SIP invitation response to the first terminal, the first SIP invitation response carrying a first SDP response, the first SDP response including first information, the first information being used to instruct the first terminal to convert video data collected during the video call into video semantic features; receiving the video semantic features sent by the first terminal, and audio data collected by the first terminal during the video call; performing timestamp alignment based on the video semantic features and audio data, and generating first video call data based on the timestamp-aligned video semantic features and audio data, and sending it to the second terminal.

[0011] This application breaks through the traditional loosely coupled pipeline limitation where the control layer is only responsible for session establishment and cannot control the media generation method by embedding the control information of generative communication into the existing SIP / SDP signaling system of IMS, thereby realizing generative communication in operator networks. Specifically, when the first terminal initiates a video call, after receiving the SIP invitation request, the IMS network element confirms that the user has activated the semantic service function and has the ability to extract video semantics based on the user's subscription data. Then, in the returned SIP invitation response, it uses the SDP response to carry the first information, explicitly instructing the first terminal to convert the raw video data collected in the video call into video semantic features. Thus, the SIP signaling, which originally could not carry the communication intent, is given the ability to dynamically adjust the media generation strategy, transforming the user's privacy protection needs such as "not exposing the real picture" into precise media layer control instructions, solving the problem that the control layer cannot drive generative strategies. Subsequently, the network side receives the video semantic features and original audio data uploaded by the first terminal, aligns them with timestamps, and generates the first video call data based on the aligned video semantic features and audio data, which is then sent to the second terminal. This deeply integrates semantic extraction at the terminal and semantic reconstruction at the network into the existing IMS session establishment and media transmission process, so that semantic communication is no longer limited to a purely wireless abstract link, but relies on standard SIP signaling interaction and service triggering mechanisms to complete capability negotiation, policy issuance, and media generation. This eliminates the disconnect between semantic communication and the existing telecommunications network, significantly reduces bandwidth pressure in weak network conditions, achieves privacy shielding of the user's real environment, and completes a full carrier-grade generative session management. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0013] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating the first embodiment of the communication method of this application. Figure 2 This is a flowchart illustrating the second embodiment of the communication method of this application. Figure 3 This is a flowchart illustrating the third embodiment of the communication method of this application. Figure 4 This is a flowchart illustrating the fourth embodiment of the communication method of this application. Figure 5 This is a schematic diagram of the framework of a communication method in one embodiment of this application; Figure 6 This is a service flow diagram of 5G new call real-time semantic communication in a communication method according to an embodiment of this application; Figure 7 This is a service flow diagram of 5G message asynchronous semantic communication in a communication method according to an embodiment of this application; Figure 8 This is a schematic diagram of the IMS network element structure of the hardware operating environment involved in the communication method in the embodiments of this application.

[0015] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0017] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0018] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments / implementations.

[0019] The following are reference definitions of some of the technical terms involved in the technical solution of this application: IMS: A core architecture for providing multimedia services over IP networks. It adopts a layered architecture design, including a control layer, a service layer, and a media layer. It implements session control through the standard SIP protocol and supports various IP multimedia services such as video calls, instant messaging, and voice calls. It is a key technology platform for realizing converged communications in operator networks.

[0020] IMS network element: A functional entity in the IMS network architecture.

[0021] SIP: An application layer control protocol used to create, modify, and terminate multimedia sessions. It uses a text-based message exchange mechanism, supports a request-response transaction model, and defines methods such as INVITE, ACK, BYE, CANCEL, REGISTER, and OPTIONS. It identifies users and resources through SIP URIs (Uniform Resource Identifiers) and is the core protocol for implementing session control in IMS networks.

[0022] SIP Request: A request message initiated by the client in the SIP protocol to request a specific operation or service from the server. It consists of three parts: a request line, a header, and a body. The request line contains the method name, request URI, and protocol version. The header contains control information such as routing, authentication, and session parameters. The body may carry session description information. Common SIP requests include SIP INVITE, SIP ACK, SIP BYE, SIP CANCEL, SIP REGISTER, and SIPOPTIONS requests.

[0023] SIP Response: A response message returned by the server in the SIP protocol to answer SIP requests. It consists of three parts: a status line, a header, and a body. The status line contains the protocol version, status code, and reason phrase. The header contains session parameters, routing information, etc., and the body may carry session description information.

[0024] SIP Invitation Request: A request message initiated using the INVITE method in the SIP protocol to invite users to participate in a multimedia session. It is the core signaling for establishing a session and includes information such as the identifier of the session initiator, the identifier of the called party, and session parameters. The message body usually carries an SDP proposal to negotiate the media type, codec, transmission parameters, etc. When the called party accepts the invitation, it will return a 200 OK response to complete the session establishment.

[0025] SIP Invitation Response: The response message returned in response to a SIP invitation request, including provisional and final responses. The message body usually carries an SDP response to confirm the media parameters of the session and complete the negotiation of media capabilities. It is a key signaling interaction in the session establishment process.

[0026] SIP Message Request: A request message initiated using the MESSAGE method in the SIP protocol. It is used to transmit instant messaging content in an established session or independently. The message body contains instant messaging data such as text, pictures, and emoticons. It supports both in-session and out-of-session message modes and is the core signaling for implementing instant messaging services in the IMS network.

[0027] SIP message response: The response message returned in response to a SIP message request, usually a 200 OK response, indicating that the message has been successfully received and processed. If the message cannot be delivered or processing fails, a corresponding error response is returned to indicate the status and result of message transmission.

[0028] SDP: A text-based protocol used to describe multimedia session parameters. It defines information such as the media type, codec, transport address, port, and bandwidth of the session. It organizes data in the form of attribute-value pairs and includes two parts: session-level description and media-level description. It is the standard carrier for media capability negotiation in the SIP protocol.

[0029] SDP Offer: An SDP description sent by the session initiator during session establishment. It includes capability information such as the media types, codec list, and transmission parameters supported by the initiator. It is used to demonstrate its own media capabilities to the called party and request the called party to respond according to these capabilities. It is the starting point for media negotiation.

[0030] SDP Answer: An SDP description returned by the called party during session establishment. It contains confirmation information such as the media type, codec, and transmission parameters selected by the called party from the SDP proposal. It is used to confirm the final media configuration adopted with the initiator, complete the negotiation of media capabilities, and is the endpoint of media negotiation.

[0031] A media line is a media description line in an SDP description that begins with "m=". It is used to define a media stream in a session and includes information such as media type, transport port, transport protocol, and codec list. One media line corresponds to one media type. An SDP description can contain multiple media lines, which describe different types of media streams such as audio, video, and data.

[0032] Media line direction refers to the direction of media stream transmission defined by attributes such as "a=sendrecv" (bidirectional transmission), "a=sendonly" (send only), "a=recvonly" (receive only), and "a=inactive" (inactive). It is used to indicate whether the media stream is bidirectional, send only, receive only, or inactive, and controls the transmission behavior of the media stream between the two parties in the session. It is an important parameter in media negotiation.

[0033] A video media line refers to a media line with the media type "video". It describes the transmission parameters of the video stream, including information such as video codec, resolution, frame rate, and bandwidth. It defines the encoding format and transmission configuration of the video stream in a video call.

[0034] The audio media line refers to the media line with the media type "audio". It is used to describe the transmission parameters of the audio stream, including information such as audio codec, sampling rate, number of channels, and bandwidth. It defines the encoding format and transmission configuration of the audio stream in video calls.

[0035] A data media line refers to a media line with the media type "application". It describes the transmission parameters of the data stream, including information such as data transmission protocol and application type. It supports the transmission of data content such as files, screen sharing, and whiteboards in a session and is the carrier of data transmission in a multimedia session.

[0036] Please refer to Figure 1 This application proposes a communication method according to a first embodiment.

[0037] In this embodiment, the communication method includes steps S100~S400: Step S100: Receive a first SIP invitation request from the first terminal. The first SIP invitation request is a request from the first terminal to initiate a video call to the second terminal.

[0038] In this embodiment, the first terminal can be a user device with video calling capabilities, such as a smartphone, tablet, laptop, or vehicle-mounted terminal, and the second terminal can also be any of the aforementioned user devices. After the user triggers a video call, the first terminal generates a first SIP invitation request and sends it to the IMS core network via the IMS access network.

[0039] The first SIP invitation request uses the INVITE method. The request line contains the SIP URI of the second terminal as the request target. The message header contains standard fields for identifying the session initiator, the called party, the session unique identifier, the request sequence number, and the routing path. The message body carries an SDP proposal for describing the media capabilities supported by the first terminal.

[0040] It should be noted that, in this embodiment, when the first terminal generates the first SIP invitation request, it can carry capability indication information in the session-level or media-level attributes of the SDP proposal to indicate to the IMS network element that it has video semantic extraction capabilities.

[0041] Optionally, the first terminal may add a custom attribute line to the SDP proposal, or add a semantic capability identifier to the user agent header field. Additionally, the first terminal may carry capability indication information in the assertion identity header field or other extended header fields of the SIP request.

[0042] In this embodiment, after receiving the first SIP invitation request, the IMS network element can parse and verify it. Optionally, the IMS network element can verify the integrity and legality of the SIP request, check whether the request header field conforms to the SIP protocol specification, verify whether the SDP proposal format is correct, and check whether parameters such as media type and codec are within the supported range. In addition, the IMS network element can also determine the subsequent signaling forwarding path based on the routing information in the request.

[0043] This embodiment achieves accurate capture and parsing of video call requests initiated by the first terminal through a standard SIP signaling reception mechanism. This request reception method based on the standard SIP protocol ensures full compatibility with existing IMS networks, requiring no large-scale modifications to the terminal or network. Furthermore, by supporting multiple capability indication methods, this embodiment provides a flexible information source for subsequent semantic service function judgment. Capability information can be transmitted through SDP extended attributes or through SIP header field extensions, offering good implementation flexibility and deployment adaptability.

[0044] In step S200, in response to the first SIP invitation request, it is determined that the first terminal has enabled the semantic service function and has the video semantic extraction capability. A first SIP invitation response is sent to the first terminal. The first SIP invitation response carries a first SDP response. The first SDP response includes first information. The first information is used to instruct the first terminal to convert the video data collected in the video call into video semantic features.

[0045] In this embodiment, after receiving the first SIP invitation request, the IMS network element needs to determine whether the first terminal has enabled semantic service functions and whether it has video semantic extraction capabilities.

[0046] In one example, an IMS network element can query the home subscriber server or user data repository to obtain the user subscription data of the first terminal. The user subscription data contains the user's service subscription information, and the IMS network element can check whether it contains subscription records for semantic service functions. If the query results show that the first terminal has subscribed to semantic service functions, it further confirms that it possesses video semantic extraction capabilities.

[0047] In one example, the IMS network element can determine that the first terminal has video semantic extraction capabilities based on the capability indication information carried in the first SIP invitation request. As mentioned before, the first terminal can carry capability indication information in the SDP proposal or SIP header field to indicate that it has video semantic extraction capabilities. After parsing this information, the IMS network element can directly confirm that the first terminal has video semantic extraction capabilities. This method does not require querying an external database, has a faster processing speed, and is suitable for scenarios with high real-time requirements.

[0048] It should be noted that, in this embodiment, video semantic extraction capability refers to the terminal's ability to convert the collected raw video data into video semantic features. Video semantic features are feature vectors or feature descriptors that can characterize the semantic information of video content. Compared with the original video data, the data volume of video semantic features is significantly reduced, while retaining the key semantic information of the video content, facilitating transmission over the network and reconstruction at the receiving end.

[0049] In this embodiment, after determining that the first terminal has enabled semantic service function and has video semantic extraction capability, the IMS network element sends a first SIP invitation response to the first terminal. The message body of the first SIP invitation response carries a first SDP response. The first SDP response contains first information, which is used to instruct the first terminal to convert the video data collected in the video call into video semantic features.

[0050] In this embodiment, there are multiple ways to carry the first information.

[0051] In one alternative implementation, the initial information can be carried as a custom attribute line in the SDP response. For example, attribute lines can be added at the session level or the video media line level. "a=x-semantic-mode:extract"; Here, "x-semantic-mode" is a custom attribute name, and "extract" indicates the extraction mode, instructing the terminal to perform video semantic extraction.

[0052] In another alternative implementation, the first information can be carried as codec parameters of the video media line in the SDP response. For example, a specific codec identifier can be added to the codec list of the video media line, and parameters for instructing the terminal to perform video semantic extraction can be defined in the attribute line of that codec, thereby transmitting video semantic extraction instructions through a codec negotiation mechanism.

[0053] In another alternative implementation, the first piece of information can be carried as a media line direction attribute in the SDP response. For example, the direction of the video media line can be set to "a=sendonly", while specifying the video semantic extraction requirements in the additional attributes.

[0054] In this embodiment, in addition to the first information indicating video semantic extraction, the first SDP response may also include other auxiliary information. For example, it may specify the format requirements of the video semantic features (such as feature dimensions and data types), extraction frequency (such as how many frames per second), and quality requirements (such as feature accuracy level). This auxiliary information can help the first terminal perform video semantic extraction operations more accurately.

[0055] This embodiment breaks the traditional loosely coupled pipeline limitation where the control layer is only responsible for session establishment and cannot control the media generation method by embedding the control information of generative communication into the existing SIP / SDP signaling system of IMS. Specifically, the IMS network element intelligently determines whether to enable semantic service functions based on user subscription data and / or terminal capability indications, and sends precise media generation policy instructions to the terminal through the first information in the SDP response, giving SIP signaling, which originally could not carry communication intent, the ability to dynamically adjust the media generation policy. This control information transmission mechanism based on the standard SIP / SDP protocol not only ensures full compatibility with the existing IMS network, but also realizes the precise driving of the media layer generation method by the control layer, solving the fundamental problem that the control layer cannot drive generative policies. At the same time, by supporting multiple ways of carrying the first information and rich auxiliary parameter configurations, this embodiment has good flexibility and scalability, and can adapt to different business scenarios and technology evolution needs.

[0056] Step S300: Receive video semantic features sent by the first terminal, and audio data collected by the first terminal during the video call.

[0057] In this embodiment, after receiving the first SIP invitation response and parsing the first information, the first terminal converts the video data it collected during the video call into video semantic features according to the instructions of the first information, while continuing to collect audio data, and sends the video semantic features and audio data to the IMS network element.

[0058] In this embodiment, the generation of video semantic features can employ various technical solutions. Optionally, the first terminal can use a deep learning-based semantic extraction model, such as a convolutional neural network, recurrent neural network, or Transformer architecture model, to extract features from the acquired video frames and generate compact semantic feature vectors. These semantic feature vectors can capture key semantic information of the video content, such as character poses, expressions, actions, and scene layouts, while significantly reducing the amount of data compared to the original video frames.

[0059] In this embodiment, there are multiple options for the transmission method of video semantic features.

[0060] Optionally, the first terminal can encapsulate video semantic features in RTP (Real-time Transport Protocol) packets and transmit them via the same RTP stream as the audio data, or via a separate RTP stream. The video semantic features and audio data can be distinguished in the RTP header using the Payload Type field.

[0061] Optionally, the first terminal can also transmit video semantic features through a data channel established via a data media line. Specifically, the first terminal can include a data media line (media type "application") in the SDP proposal to negotiate and establish a data transmission channel, and then transmit the video semantic features through this data channel. This approach can utilize reliable transmission protocols such as SCTP (Stream Control Transmission Protocol) to ensure the integrity and reliability of the video semantic features, making it particularly suitable for scenarios with high transmission quality requirements.

[0062] It should be noted that, in this embodiment, when the first terminal sends video semantic features and audio data, it needs to add timestamp information to each data unit. The timestamp is used to identify the acquisition time or generation time of the data unit and is the basis for subsequent timestamp alignment.

[0063] In this embodiment, the audio data is processed in the same way as in traditional video calls. The first terminal collects audio signals through a microphone, encodes them using an audio codec, and then encapsulates them in RTP data packets before sending. The audio data is transmitted in its original form without semantic extraction or conversion to ensure the real-time nature and naturalness of the voice communication.

[0064] This embodiment achieves a distributed architecture where semantic extraction is performed at the terminal and semantic reconstruction is performed on the network by receiving video semantic features and audio data sent by the first terminal. This architecture fully utilizes the terminal's computing power for semantic feature extraction, while offloading the computationally intensive semantic reconstruction task to the network side, thus achieving a reasonable allocation of computing resources and significantly reducing the data transmission pressure between the first terminal and IMS network elements. Compared to raw video data, video semantic features have a significant advantage in data volume, greatly reducing bandwidth pressure and improving transmission reliability and stability in weak network environments. Furthermore, since semantic features are transmitted instead of raw video frames, the user's facial image and real environment are effectively protected, meeting the user's privacy requirements. By supporting multiple semantic feature generation schemes and transmission methods, this embodiment has good flexibility and adaptability, enabling the selection of the optimal implementation scheme based on different terminal capabilities and network conditions.

[0065] Step S400: Timestamp alignment is performed based on video semantic features and audio data, and first video call data is generated and sent to the second terminal based on the timestamp-aligned video semantic features and audio data.

[0066] In this embodiment, after receiving the video semantic features and audio data sent by the first terminal, the IMS network element first performs timestamp alignment. The purpose of timestamp alignment is to ensure that the video semantic features and audio data are synchronized in the time dimension, avoiding the problem of audio-visual asynchrony.

[0067] In this embodiment, after completing the timestamp alignment, the IMS network element generates the first video call data based on the aligned video semantic features and audio data.

[0068] In one feasible implementation, the generation process of the first video call data can be divided into two main stages: video semantic reconstruction and audio-visual synthesis.

[0069] In the video semantic reconstruction stage, IMS network elements utilize pre-trained semantic reconstruction models to convert video semantic features into visualized video frames. Optionally, generative models such as generative adversarial networks, variational autoencoders, and diffusion models can be used for semantic reconstruction. These models can generate video frames that conform to the semantic description based on the input semantic features, and can adjust the generation effect according to preset style parameters (such as cartoon style, realistic style, abstract style, etc.).

[0070] In one alternative implementation, the semantic reconstruction model can generate personalized virtual avatar videos. For example, it can generate specific virtual avatars based on user preferences or stylized virtual avatars based on user facial features, thus protecting user privacy while providing a personalized visual experience.

[0071] In another alternative implementation, the semantic reconstruction model can generate contextualized video content. For example, based on scene layout information in the video's semantic features, a corresponding virtual scene can be generated, and character actions can be overlaid onto the virtual scene to achieve a richer visual presentation.

[0072] In the audio and video synthesis stage, the IMS network element combines the reconstructed video frames with audio data to generate a complete video call data stream. Optionally, standard audio and video multiplexing technology can be used to encapsulate video frames and audio data in a unified container format, or encapsulate them separately in independent RTP streams for transmission.

[0073] In this embodiment, the generated first video call data is transmitted to the second terminal via the IMS network. The transmission process can employ the standard IMS media transmission mechanism, using network elements such as media gateways and border gateways for routing and forwarding. After receiving the first video call data, the second terminal decodes and plays it, completing the receiving end processing of the video call.

[0074] It is worth noting that in this embodiment, the entire processing flow is deeply integrated into the existing IMS session establishment and media transmission process. From SIP signaling interaction, capability negotiation, and policy issuance to media generation, transmission, and playback, all steps are completed based on the standard IMS protocol and service triggering mechanism, without requiring large-scale modifications to the existing network architecture, thus demonstrating good feasibility for deployment on the existing network.

[0075] This embodiment deeply integrates semantic extraction at the terminal and semantic reconstruction at the network architecture into the existing IMS session establishment and media transmission process through timestamp alignment and semantic reconstruction mechanisms. Timestamp alignment ensures audio-visual synchronization, providing a good user experience; semantic reconstruction converts compact semantic features into visualized video content, realizing the core function of generative communication. Compared with the purely wireless abstract link semantic communication model in academia, this embodiment relies on standard SIP signaling interaction and service triggering mechanisms to complete the entire process of capability negotiation, policy issuance, and media generation, enabling semantic communication to move beyond theoretical research and be directly deployed in operator networks, achieving deep integration of semantic communication with existing telecommunications networks. In weak network environments, because the transmitted video semantic features have a significantly reduced data volume, bandwidth pressure is greatly reduced, and communication quality is significantly improved; at the same time, since the reconstructed video content is a virtual image or scene generated based on semantic features, the user's facial image and real environment are effectively protected, meeting the user's privacy protection needs. By supporting multiple semantic reconstruction models and style configurations, this embodiment also provides a personalized visual experience, further improving user satisfaction.

[0076] This embodiment breaks through the traditional loosely coupled pipeline limitation where the control layer is only responsible for session establishment and cannot control the media generation method by embedding the control information of generative communication into the existing SIP / SDP signaling system of IMS, thereby realizing generative communication in the operator network. Specifically, when the first terminal initiates a video call, after receiving the SIP invitation request, the IMS network element confirms that the user has activated the semantic service function and has the ability to extract video semantics based on the user's subscription data. Then, in the returned SIP invitation response, it uses the SDP response to carry the first information, explicitly instructing the first terminal to convert the raw video data collected in the video call into video semantic features. Thus, the SIP signaling, which originally could not carry the communication intent, is given the ability to dynamically adjust the media generation strategy, transforming the user's privacy protection needs such as "not exposing the real picture" into precise media layer control instructions, solving the problem that the control layer cannot drive generative strategies. Subsequently, the network side receives the video semantic features and original audio data uploaded by the first terminal, aligns them with timestamps, and generates the first video call data based on the aligned video semantic features and audio data, which is then sent to the second terminal. This deeply integrates semantic extraction at the terminal and semantic reconstruction at the network into the existing IMS session establishment and media transmission process, so that semantic communication is no longer limited to a purely wireless abstract link, but relies on standard SIP signaling interaction and service triggering mechanisms to complete capability negotiation, policy issuance, and media generation. This eliminates the disconnect between semantic communication and the existing telecommunications network, significantly reduces bandwidth pressure in weak network conditions, achieves privacy shielding of the user's real environment, and completes a full carrier-grade generative session management.

[0077] In one feasible implementation, the first SDP response further includes second information, which instructs the first terminal to send video semantic features and audio data collected by the first terminal during the video call to the target address.

[0078] The above-mentioned step S300, receiving video semantic features sent by the first terminal and audio data collected by the first terminal during the video call, may include step S310: Step S310: Based on the target address, receive the video semantic features sent by the first terminal, as well as the audio data collected by the first terminal during the video call.

[0079] In this embodiment, the core function of the second information is to reconstruct the media transmission topology. In traditional video calls, the video and audio data of the first terminal are directly sent to the second terminal, forming a point-to-point dual-transmit and dual-receive topology. The IMS network only acts as a transparent conduit and cannot modify or generate the transmitted media data. However, this embodiment uses the second information to set the target address to a functional entity in the IMS network responsible for semantic reconstruction, such as a semantic reconstruction server or a specific IMS network element, thereby achieving a fundamental reconstruction of the media topology from the traditional point-to-point mode to a star-shaped generated topology.

[0080] Specifically, the second piece of information can be carried as a custom attribute line in the SDP response, for example, by adding an attribute line: "a=x-semantic-target:192.168.1.100"; Here, "192.168.1.100" represents the address of the semantic reconstruction function entity in the IMS network. The target address can be the network address of the functional module responsible for receiving and processing semantic features in the IMS network element, or it can be the address of a specially deployed semantic processing server.

[0081] It should be noted that in this embodiment, the selection of the target address can be dynamic. The IMS network element can dynamically select the optimal semantic reconstruction functional entity as the target address based on factors such as network load, semantic reconstruction resource availability, and geographical location, and inform the first terminal through the second information. This dynamic target address allocation mechanism enables generative communication to flexibly adapt to different network environments and resource conditions.

[0082] In this embodiment, since the second information sets the target address to the semantic reconstruction function entity in the IMS network, step S310 is actually executed on that semantic reconstruction function entity. When the first terminal sends video semantic features and audio data to the target address according to the instructions of the second information, the semantic reconstruction function entity receives this data at that address, rather than the second terminal receiving it directly.

[0083] In this embodiment, this target address-based reception mechanism is a key component of the generative communication architecture. The semantic reconstruction functional entity, acting as the central node in the star topology, receives semantic features and audio data from the first terminal, performs semantic reconstruction processing, and then sends the reconstructed video call data to the second terminal. This architecture allows the network side to fully control the media generation process, achieving generative capabilities that are impossible with traditional point-to-point communication.

[0084] It is important to emphasize that, in this embodiment, the semantic reconstruction function entity can be a dedicated network element in the IMS network or an extended functional module of an existing IMS network element. Regardless of the implementation method, its core function is to act as the central node of the star topology, receive uplink semantic features, perform reconstruction processing, and send down the reconstructed media data.

[0085] This implementation achieves a fundamental reconstruction of the media transmission topology by including second information in the first SDP response. In traditional video calls, the first and second terminals operate in a point-to-point, dual-transmit and dual-receive mode, and IMS network elements cannot modify the transmitted data. However, this embodiment constructs a star-shaped generative topology by setting the target address to the semantic reconstruction function entity on the network side: the first terminal sends semantic features and audio uplink to the network side, and the network side reconstructs and sends it downlink to the second terminal. This topology reconstruction allows IMS network elements to be inserted into the media transmission path between terminals to perform semantic reconstruction and generation processing on the transmitted data, thereby realizing the basic architecture of generative communication.

[0086] In one feasible implementation, the above communication method may further include step S210: In step S210, in response to the first SIP invitation request, it is determined that the first terminal has enabled semantic service function and has video semantic extraction capability, and a second SIP invitation request is also sent to the second terminal. The second SIP invitation request carries the first SDP proposal.

[0087] The first SDP proposal includes third information, which instructs the second terminal to receive first video call data from the target address during a video call.

[0088] In this embodiment, step S210 is another key step in the star topology. After the IMS network element determines that the first terminal has enabled semantic service functions and has video semantic extraction capabilities, in addition to sending a first SIP invitation response to the first terminal, it also needs to send a second SIP invitation request to the second terminal, informing the second terminal to receive the video call data reconstructed from the network side.

[0089] In this embodiment, the generation and transmission of the second SIP invitation request can refer to the standard IMS session establishment process. The difference lies in that the first SDP proposal carried in the second SIP invitation request declares the media capabilities of the IMS network element, especially the target address indicated by the third information. This ensures that after receiving the second SIP invitation request, the second terminal can establish a connection with the network-side semantic reconstruction function entity according to the indication of the third information and the media capabilities declared in the first SDP proposal, and receive the reconstructed first video call data. This connection establishment mechanism ensures the integrity of the star topology: the first terminal sends uplink data to the network side, and the second terminal receives downlink data from the network side, with the network side acting as the central node to complete the semantic reconstruction.

[0090] This implementation constructs a star topology by sending a second SIP invitation request containing third information to the second terminal. In traditional video calls, the second terminal directly receives video data from the first terminal; however, in this embodiment, the second terminal receives reconstructed video data from the network-side semantic reconstruction function entity. This architectural restructuring allows the network side to fully control the media generation process, converting the semantic features uploaded by the first terminal into video content that meets specific needs (such as virtual avatars, stylized videos, etc.), thus realizing the core capabilities of generative communication. Furthermore, by embedding the receiver configuration information into the standard SIP / SDP signaling system, this implementation maintains full compatibility with existing IMS networks, requiring no large-scale modifications to the terminal or network.

[0091] In one feasible implementation, when the first terminal has enabled the semantic service function and the second terminal has not enabled the semantic service function, the first SDP response further includes first video media line direction information for indicating that the video media line is only sent, first audio media line direction information for indicating that the audio media line is transmitted bidirectionally, and first data media line direction information for indicating that the data media line is only received.

[0092] The first SDP proposal also includes second video media line direction information for indicating bidirectional transmission of video media lines, second audio media line direction information for indicating bidirectional transmission of audio media lines, and second data media line direction information for indicating inactive data media lines.

[0093] In this embodiment, the direction of the video media line in the first SDP response is set to "a=sendonly", indicating that the IMS network element only sends and does not receive in the video media channel established with the first terminal. This means that the IMS network element is informing the first terminal: I will send you video data, but you do not need to send me video data. In a star topology of generative communication, this indicates that the network side will send video content to the first terminal through this video media channel, usually video data uploaded by the second terminal, but the first terminal does not need to upload video data to this channel.

[0094] In the first SDP response, the direction of the audio media line is set to "a=sendrecv", indicating that the IMS network element both sends and receives audio media data in the audio media channel established with the first terminal. This means that the IMS network element is informing the first terminal: You can send me audio data, and I will also send you audio data. In a star topology for generative communication, this ensures bidirectional voice communication between the first and second terminals, with audio data being bidirectionally forwarded through the IMS network element.

[0095] In the first SDP response, the direction of the data media line is set to "a=recvonly", indicating that the IMS network element only receives and does not send data media in the data media channel established with the first terminal. This means that the IMS network element is telling the first terminal: You can send data media to me, but I will not send data media to you. In a star topology of generative communication, this provides a dedicated data channel for the first terminal to upload video semantic features. The first terminal sends the video semantic features to the IMS network element for semantic reconstruction processing through this data media channel.

[0096] In the first SDP proposal, the direction of the video media line is set to "a=sendrecv", indicating that the IMS network element both sends and receives in the video media channel established with the second terminal. This means that the IMS network element informs the second terminal: you can send me video data, and I will also send you video data. In a star topology of generative communication, this indicates that the network side will send video data reconstructed based on the semantic features of the first terminal to the second terminal, and also supports the second terminal uploading its own video data.

[0097] In the first SDP proposal, the direction of the audio media line is set to "a=sendrecv", indicating that the IMS network element both sends and receives audio media data in the audio media channel established with the second terminal. This means that the IMS network element informs the second terminal: you can send audio data to me, and I will also send audio data to you. In a star topology for generative communication, this ensures bidirectional voice communication between the second terminal and the first terminal, with audio data being bidirectionally forwarded through the IMS network element.

[0098] In the first SDP proposal, the direction of the data media line is set to "a=inactive," indicating that the IMS network element neither sends nor receives data in the data media channel established with the second terminal. This means that the IMS network element informs the second terminal: a data media channel does not need to be established between us. In a star topology for generative communication, this avoids unnecessary data media resource consumption on the second terminal side, because the second terminal does not need to upload video semantic features.

[0099] This implementation achieves an accurate description of the media capabilities of IMS network elements in a star-shaped topology by precisely configuring the direction information of each media line in the first SDP response and the first SDP proposal. For the first terminal side, the video media line is set to sendonly, indicating that the network side sends video data from the second terminal to the first terminal; the audio media line is set to sendrecv (bidirectional transmission) to maintain the bidirectional nature of voice communication; and the data media line is set to receive only, providing a dedicated channel for the first terminal to upload video semantic features. For the second terminal side, the video media line is set to sendrecv (bidirectional transmission), indicating that the network side sends the reconstructed video to the second terminal and also supports the second terminal uploading video; the audio media line is set to sendrecv (bidirectional transmission) to maintain the bidirectional nature of voice communication; and the data media line is set to inactive to avoid unnecessary resource consumption. This media routing configuration accurately reflects the star topology of generative communication: the first terminal sends video semantic features uplink through the data media channel, transmits audio bidirectionally through the audio media channel, and receives video data from the second terminal on the network side; the second terminal receives the reconstructed video data through the video media channel and transmits audio bidirectionally through the audio media channel; the network side, as the central node, receives the video semantic features from the first terminal, performs semantic reconstruction, and then sends the reconstructed video data downlink to the second terminal, while simultaneously forwarding audio data bidirectionally. This topology reconstruction allows the IMS network to be inserted into the media transmission path between terminals, performing semantic reconstruction and generation processing on the transmitted data, thereby realizing the core capabilities of generative communication.

[0100] Please refer to Figure 2 This application proposes a communication method according to a second embodiment.

[0101] In the second embodiment of this application, the same or similar content as in the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0102] In this embodiment, the video semantic features are structural semantic features that characterize the activity information of people in the video data. The activity information of people includes at least one of the following: person's posture, person's actions, and person's facial expressions.

[0103] It should be noted that, in this embodiment, structural semantic features refer to feature representations that can characterize a person's activity information in a structured manner. Unlike traditional global feature vectors, structural semantic features adopt a hierarchical and modular representation, independently encoding and describing different dimensions of a person's activity.

[0104] In this embodiment, the structural semantic features of a person's posture may include information such as the coordinates of the person's skeletal joints, limb angles, and body orientation. Optionally, a human posture estimation algorithm can be used to extract the person's 2D or 3D skeletal joint information, and these joint coordinates can be used as part of the structural semantic features. This joint information can accurately describe the person's body posture, providing accurate motion parameters for subsequent digital human image driving.

[0105] In this embodiment, the structural semantic features of a person's actions may include information such as action category labels, action intensity, action duration, and action trajectory. Optionally, an action recognition algorithm can be used to analyze the video sequence to identify the specific actions performed by the person, such as waving, nodding, or turning around, and the temporal features of the actions can be encoded into structural semantic features. These action features can capture the dynamic behavior of the person, enabling the digital avatar to accurately reproduce the user's actions.

[0106] In this embodiment, the structural semantic features of a person's facial expressions may include information such as facial key point coordinates, expression category labels, and expression intensity. Optionally, a facial expression recognition algorithm can be used to extract facial feature points, identify the person's expression category, such as happiness, sadness, or surprise, and encode the expression intensity parameter as a structural semantic feature. These expression features can capture the person's emotional state, enabling the digital avatar to vividly represent the user's emotional changes.

[0107] It should be noted that, in this embodiment, the components of the structural semantic features can be extracted and encoded independently, or they can be fused together. Optionally, the features of a person's posture, actions, and expressions can be spatiotemporally aligned and fused to generate a unified structural semantic feature representation, thereby improving the integrity and consistency of the features.

[0108] In this embodiment, when generating video semantic features, the first terminal can choose to extract some or all of the person's activity information based on its own computing power and business needs. For example, in scenarios with limited computing resources, only the person's posture information can be extracted; in scenarios with high requirements for interactive experience, the person's posture, actions, and facial expressions can be extracted simultaneously. This flexible feature extraction strategy can adapt to different terminal capabilities and application scenarios.

[0109] This embodiment defines video semantic features as structural semantic features that characterize a person's active information, enabling refined modeling and transmission of dynamic behavior. Compared to traditional global feature representation, structural semantic features employ a hierarchical and modular encoding approach, independently describing a person's posture, actions, and expressions. This not only improves the expressive power and accuracy of the features but also provides rich control parameters for subsequent digital human avatar driving. This structured feature representation allows the digital human avatar to accurately reproduce the user's body posture, dynamic actions, and emotional expressions, significantly enhancing the naturalness and realism of the generated video and providing users with a more immersive video call experience.

[0110] The step S400 above, which generates the first video call data based on the timestamp-aligned video semantic features and audio data, may include steps S410 to S420: Step S410: Obtain the video reconstruction auxiliary parameters corresponding to the first terminal, wherein the video reconstruction auxiliary parameters include the digital human image.

[0111] In this embodiment, before generating the first video call data, the IMS network element needs to obtain the video reconstruction auxiliary parameters corresponding to the first terminal. The video reconstruction auxiliary parameters are auxiliary information used to guide the semantic reconstruction process, among which the core parameter is the digital human image.

[0112] In this embodiment, a digital human avatar refers to a virtual character used to replace a user's real-life image. The digital human avatar can be a preset, generic avatar or a user-defined, personalized avatar. Optionally, the digital human avatar may include the character's physical features (such as hairstyle, hair color, facial contours, clothing style, etc.), body features (such as height, body shape, proportions, etc.), and style features (such as cartoon style, realistic style, abstract style, etc.).

[0113] In this embodiment, there are multiple ways to obtain the digital human image.

[0114] In one example, an IMS network element can query the digital avatar configuration corresponding to the first terminal from a user data repository or home subscriber server. Before using semantic communication services, users can pre-configure and store their digital avatar preferences on the terminal or network side. When processing a video call request, the IMS network element can retrieve the corresponding digital avatar configuration from the database based on the user identifier of the first terminal.

[0115] In one example, the IMS network element can dynamically obtain digital avatar information during SIP signaling interactions. For instance, the first terminal can carry the digital avatar's identifier or configuration parameters in a SIP invitation request or subsequent SIP message. After parsing this information, the IMS network element can directly obtain the digital avatar configuration. This approach allows users to dynamically switch their digital avatar during a call, improving flexibility.

[0116] In one example, IMS network elements can automatically select a digital avatar based on business policies or network conditions. For instance, in bandwidth-constrained scenarios, a simplified digital avatar with a smaller data volume can be selected; in scenarios requiring higher visual quality, a high-precision digital avatar with a larger data volume can be selected. This adaptive selection mechanism can optimize resource utilization and user experience based on actual conditions.

[0117] In this embodiment, in addition to the digital human image, the video reconstruction auxiliary parameters may also include other auxiliary information. For example, they may include scene background parameters (used to specify the background environment of the virtual scene), lighting parameters (used to control the lighting effects during rendering), camera parameters (used to control the rendering viewpoint and depth of field effects), etc. These auxiliary parameters can further improve the visual quality and personalization of the generated video.

[0118] This embodiment obtains auxiliary parameters for video reconstruction corresponding to the first terminal, especially the digital human avatar, providing personalized visual templates and rendering guidance for the semantic reconstruction process. As a representative of the user in virtual space, the digital human avatar not only protects the user's real privacy but also offers rich personalization options, satisfying the user's diverse needs for avatar expression. Simultaneously, by combining auxiliary parameters such as scene background, lighting, and camera, this embodiment can generate richer and more realistic virtual video content, significantly improving the visual experience and immersion of video calls.

[0119] Step S420: Based on the video reconstruction auxiliary parameters, as well as the video semantic features and audio data after timestamp alignment, generate the first video call data.

[0120] Among them, video semantic features are used to drive the digital human image to be displayed actively, and the active display includes the display of at least one of the human posture, human action and human expression.

[0121] In this embodiment, after obtaining the video reconstruction auxiliary parameters, the IMS network element generates the first video call data based on the video reconstruction auxiliary parameters, the timestamp-aligned video semantic features, and the audio data. The core of this generation process is to use video semantic features to drive the digital human image to be actively displayed.

[0122] In this embodiment, the driving process of the digital human image can be divided into multiple levels, corresponding to the driving of the human posture, human actions and human expressions respectively.

[0123] At the character pose driving level, IMS network elements adjust the skeletal pose of the digital human avatar based on character pose information (such as skeletal joint coordinates and limb angles) from the video semantic features. Optionally, skeletal animation technology can be used to map the extracted character pose parameters onto the skeletal system of the digital human avatar, ensuring that the digital human avatar's pose matches the user's real pose. For example, when the user raises their arm, the digital human avatar will also raise their arm accordingly; when the user turns around, the digital human avatar will also turn around accordingly.

[0124] At the character action driving level, IMS network elements control the digital human avatar to perform corresponding actions based on character action information (such as action category, action intensity, and action trajectory) from the video semantic features. Optionally, motion capture redirection technology can be used to map the identified user actions to the digital human avatar's action library, driving the digital human avatar to perform the corresponding actions. For example, when a user waves, the digital human avatar will perform a predefined waving action; when a user nods, the digital human avatar will perform a nodding action.

[0125] At the facial expression driving level, IMS network elements adjust the digital human's facial expressions based on facial expression information (such as facial key point coordinates, expression category, and expression intensity) from the video's semantic features. Optionally, facial animation technology can be used to map the extracted expression parameters onto the digital human's facial controller, ensuring that the digital human's expressions are synchronized with the user's real expressions. For example, when the user smiles, the digital human will also display a smiling expression; when the user frowns, the digital human will also display a corresponding frowning expression.

[0126] It should be noted that in this embodiment, the character's posture, actions, and expressions can be driven independently or in combination.

[0127] In this embodiment, the rendering process of the digital human image can be optimized based on other parameters in the video reconstruction auxiliary parameters. For example, a corresponding virtual scene can be rendered based on scene background parameters, and the driven digital human image can be superimposed on the scene; the lighting effect during rendering can be adjusted based on lighting parameters to match the lighting conditions of the scene with the digital human image; and the rendering perspective and depth of field effect can be controlled based on camera parameters to provide a more natural visual experience.

[0128] In this embodiment, the audio-visual synthesis stage combines the rendered digital human avatar video with audio data. Optionally, lip-sync technology can be used to fine-tune the avatar's mouth movements based on the speech content in the audio data, ensuring that the mouth movements are synchronized with the speech content, further enhancing the realism and naturalness of the video.

[0129] It should be noted that in this embodiment, the entire generation process can employ real-time rendering technology to ensure the real-time nature and smoothness of the video call. Optionally, GPU (Graphics Processing Unit) accelerated rendering, distributed rendering, and other technologies can be used to improve rendering efficiency and reduce processing latency. Simultaneously, rendering quality and frame rate can be dynamically adjusted based on network bandwidth and terminal capabilities to optimize resource utilization while ensuring a good user experience.

[0130] This embodiment generates first video call data based on video reconstruction auxiliary parameters, timestamp-aligned video semantic features, and audio data, achieving precise driving and active display of the digital human avatar. Specifically, it utilizes the character's posture, actions, and facial expressions from structural semantic features to drive the skeletal posture, action execution, and facial expressions of the digital human avatar, enabling it to accurately reproduce the user's dynamic behavior and emotional state. This multi-layered and refined driving mechanism not only ensures the naturalness and realism of the digital human avatar but also achieves effective substitution for the user's real image and privacy protection. By combining auxiliary parameters such as scene background, lighting, and camera, this embodiment can generate richer and more realistic virtual video content, providing users with a personalized visual experience. Simultaneously, by employing real-time rendering technology and optimization strategies, this embodiment can achieve low-latency, high-smoothness video calls while ensuring video quality, meeting the stringent requirements of real-time communication.

[0131] This embodiment further deepens the technical connotation of generative communication by defining video semantic features as structural semantic features representing the active information of a person and displaying them actively based on a digital human image. Compared with the general semantic reconstruction in the first embodiment, this embodiment adopts structured feature representation and a refined driving mechanism to achieve accurate modeling and reproduction of the dynamic behavior of a person. As a representative of the user in virtual space, the digital human image not only effectively protects the user's real privacy but also provides rich personalized choices and expression methods. Through multi-level driving strategies and real-time rendering technology, this embodiment can achieve low-latency and high-smoothness video calls while ensuring video quality, meeting the stringent requirements of real-time communication. This generative communication scheme based on structural semantic features and digital human images provides a more practical and efficient implementation path for semantic communication in operator networks, and has significant application value and promotion prospects.

[0132] In one feasible implementation, step S420 above, based on video reconstruction auxiliary parameters and timestamp-aligned video semantic features and audio data, generates the first video call data, and may include steps S421 to S423: Step S421: Obtain the current network quality parameters of the video call, wherein the network quality parameters include transmission delay and / or packet loss rate.

[0133] In this embodiment, before generating the first video call data, the IMS network element needs to obtain the current network quality parameters of the video call. Network quality parameters are key indicators reflecting the current communication link transmission performance, directly affecting the data transmission efficiency and user experience of the video call.

[0134] In this embodiment, network quality parameters may include at least one of transmission delay and packet loss rate. Transmission delay refers to the time required for data to travel from the sender to the receiver, usually measured in milliseconds; packet loss rate refers to the proportion of data packets lost during data transmission out of the total number of transmitted data packets, usually expressed as a percentage.

[0135] This embodiment obtains real-time network quality parameters for video calls, providing accurate decision-making basis for subsequent video rendering strategy adjustments. Based on key indicators such as transmission latency and packet loss rate, the IMS network element can dynamically perceive changes in network conditions and adjust video generation strategies in a timely manner, ensuring a stable and smooth video call experience under different network conditions.

[0136] Step S422: Determine video rendering strategy parameters based on network quality parameters. The video rendering strategy parameters include video resolution and / or the parameter magnitude of the generation model that generates the first video call data. The higher the parameter magnitude, the greater the computational resource requirement of the generation model.

[0137] In this embodiment, after obtaining the network quality parameters, the IMS network element needs to determine the video rendering strategy parameters based on these parameters. The video rendering strategy parameters are key configurations guiding the video generation process, directly affecting the quality, computational complexity, and transmission efficiency of the generated video.

[0138] In this embodiment, the video rendering strategy parameters may include at least one of the video resolution and the parameter magnitude of the generation model.

[0139] Regarding video resolution, different video resolutions can be selected depending on the network quality parameters. Optionally, when the network quality is good, such as low transmission latency and low packet loss rate, a higher video resolution can be selected to provide a clearer and more detailed visual experience; when the network quality is poor, such as high transmission latency and high packet loss rate, a lower video resolution can be selected to reduce the amount of data transmitted and improve the reliability and stability of transmission.

[0140] Regarding the parameter scale of the generative model, different parameter scales can be selected based on the different network quality parameters. Optionally, the parameter scale refers to the model size of the generative model, usually measured by the number of model parameters. The higher the parameter scale, the better the expressive power and generation quality of the generative model, but at the same time, the greater the demand for computing resources and the longer the inference latency.

[0141] In one example, when the network quality is good, a generative model with a high number of parameters can be selected, such as a large GAN (Generative Adversarial Network) model or a diffusion model with a high number of parameters, to generate higher quality and more realistic video content; when the network quality is poor, a generative model with a low number of parameters can be selected, such as a lightweight GAN model or a simplified diffusion model, to reduce computational complexity, shorten inference time, and ensure the real-time nature of video generation.

[0142] It's easy to understand that, in addition to video resolution and the magnitude of the generation model's parameters, rendering quality can also be included as a video rendering strategy parameter.

[0143] It should be noted that, in this embodiment, the adjustment of video rendering strategy parameters can be gradual or abrupt. Optionally, to avoid sudden changes in video quality affecting user experience, a gradual adjustment strategy can be adopted, gradually changing the video resolution or the magnitude of model parameters; in scenarios requiring rapid response to changes in network conditions, an abrupt adjustment strategy can also be adopted, directly switching to the target parameter configuration.

[0144] This embodiment achieves adaptive optimization of the video generation process by dynamically determining video rendering strategy parameters based on network quality parameters. Specifically, based on network quality indicators such as transmission latency and packet loss rate, it intelligently selects appropriate video resolution and generation model parameter levels. This provides a high-quality video experience under good network conditions and prioritizes video smoothness and stability under poor network conditions. This adaptive rendering strategy adjustment mechanism not only fully utilizes network resources but also effectively copes with network fluctuations, significantly improving the robustness of video calls and the user experience.

[0145] Step S423: Generate the first video call data based on the video rendering strategy parameters, video reconstruction auxiliary parameters, and video semantic features and audio data after timestamp alignment.

[0146] In this embodiment, after determining the video rendering strategy parameters, the IMS network element generates the first video call data based on the video rendering strategy parameters, video reconstruction auxiliary parameters, timestamp-aligned video semantic features, and audio data. This generation process integrates the video rendering strategy parameters into each stage of video generation, achieving fine-grained control over the generation process.

[0147] In the video resolution control stage, the IMS network element adjusts the output size of the generated video according to the determined video resolution parameters. Optionally, in the digital human image rendering stage, the resolution of the rendering target can be set so that the generated video frames meet the specified resolution requirements. For example, when the video resolution parameter is 1080p, the rendered output video frame size is 1920×1080 pixels; when the video resolution parameter is 720p, the rendered output video frame size is 1280×720 pixels.

[0148] During the generative model selection phase, the IMS network element selects the appropriate semantic reconstruction model for video generation based on the determined parameter level of the generative model. Optionally, multiple generative models with different parameter levels can be pre-deployed, and models can be dynamically loaded and switched according to parameter level requirements. For example, when the parameter level is high, a large GAN model or a diffusion model with a large number of parameters is loaded; when the parameter level is low, a lightweight generative model or a simplified model is loaded.

[0149] In the rendering quality optimization phase, IMS network elements can adjust other quality-related parameters during the rendering process based on video rendering strategy parameters. Optionally, rendering parameters such as texture quality, shadow effects, and anti-aliasing levels can be adjusted to further optimize computational efficiency and resource utilization while ensuring basic visual effects. For example, when network quality is poor, texture quality and shadow effects can be reduced to decrease GPU computational load; when network quality is good, high-quality rendering effects can be enabled to improve the visual experience.

[0150] It should be noted that in this embodiment, the rendering strategy parameters can be adjusted in real time during the call based on the dynamic changes in network quality, thereby achieving fine adaptive control.

[0151] In this embodiment, the generated first video call data is transmitted to the second terminal via the IMS network. The transmission process can employ the standard IMS media transmission mechanism, dynamically adjusting the transmission strategy (such as forward error correction, retransmission mechanism, etc.) based on the current network quality parameters to further improve the reliability and efficiency of transmission.

[0152] This embodiment generates first video call data based on video rendering strategy parameters, video reconstruction auxiliary parameters, and video semantic features and audio data aligned with timestamps, achieving fine-grained control and adaptive optimization of the video generation process. Specifically, rendering strategy parameters are integrated into each stage of call video reconstruction, enabling the video generation process to dynamically adjust according to network conditions. This maximizes network resource utilization while ensuring video quality, guaranteeing the smoothness and stability of the video call. This network quality-aware adaptive video generation mechanism not only effectively addresses network fluctuations but also provides the optimal video experience under different network conditions, significantly improving the practicality and reliability of generative communication.

[0153] In one feasible implementation, after obtaining the current network quality parameters of the video call in step S421, the communication method may further include steps S424-S425: Step S424: When the network quality parameters indicate that the current network quality of the video call is the first network quality, a first frequency instruction is sent to the first terminal. The first frequency instruction is used to instruct the first terminal to update the video semantic features at the first frequency.

[0154] Step S425: When the network quality parameters indicate that the current network quality of the video call is the second network quality, a second frequency instruction is sent to the first terminal. The second frequency instruction is used to instruct the first terminal to update the video semantic features at the second frequency.

[0155] Among them, the quality of the first network is lower than that of the second network, and the first frequency is lower than that of the second frequency.

[0156] In this embodiment, network quality parameters may include transmission delay and / or packet loss rate. A first network quality can be defined as a network condition where the transmission delay exceeds a preset transmission delay or the packet loss rate exceeds a preset packet loss rate; a second network quality can be defined as a network condition where the transmission delay is lower than a preset transmission delay and the packet loss rate is lower than a preset packet loss rate.

[0157] In this embodiment, the frequency command can be sent to the first terminal via SIP signaling messages, RTCP (Real-time Transport Control Protocol) feedback messages, or a dedicated control channel. After receiving the frequency command, the first terminal adjusts the update frequency of the video semantic features according to the command requirements, for example, by adjusting the frame rate of video acquisition or the sampling interval of semantic feature extraction.

[0158] This embodiment achieves adaptive control of the video semantic feature update frequency by sending frequency commands to the first terminal based on network quality parameters. Specifically, when the network quality is poor, the first terminal is instructed to update the video semantic features at a lower frequency, effectively reducing data transmission volume and network load, and improving communication reliability and stability in weak network environments. When the network quality is good, the first terminal is instructed to update the video semantic features at a higher frequency, making full use of network resources and providing a smoother, more real-time video call experience. This frequency adaptive mechanism based on network quality awareness, combined with the video rendering strategy adaptive adjustment mechanism in steps S421-S423, forms a complete end-to-end network adaptive system: from the terminal data acquisition frequency to the server-side video rendering strategy, intelligent optimization of the entire link is achieved. By supporting multi-level network quality division and flexible frequency configuration, this embodiment can dynamically adjust the system working mode according to different network conditions, providing optimal video call services in various network environments, significantly improving the practicality and user experience of generative communication.

[0159] Please refer to Figure 3 This application proposes a communication method according to a third embodiment.

[0160] In the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0161] In this embodiment, before obtaining the video reconstruction auxiliary parameters corresponding to the first terminal in step S410, steps S401-S402 may be included: Step S401: If the first SIP invitation request carries a unified semantic identifier, obtain the video reconstruction auxiliary parameters associated with the unified semantic identifier from the preset semantic communication configuration library.

[0162] In this embodiment, the unified semantic identifier is an identifier used to uniquely identify semantic communication between two communicating parties. It can be a globally unique string, number, or UUID (Universally Unique Identifier). The semantic communication configuration library is a database used to store the association between the unified semantic identifier and video reconstruction auxiliary parameters.

[0163] In this embodiment, the semantic communication configuration library can adopt a key-value pair storage structure, where the key is a unified semantic identifier and the value is the corresponding video reconstruction auxiliary parameter. The video reconstruction auxiliary parameters may include digital human image configuration, scene background parameters, lighting parameters, camera parameters, etc. The semantic communication configuration library can be deployed locally on the IMS network element or in a distributed caching system to support high-concurrency access and fast querying.

[0164] In this embodiment, when the IMS network element receives the first SIP invitation request, it first parses the request message header to check whether it contains a unified semantic identifier. If a unified semantic identifier is detected, it directly queries the semantic communication configuration library based on the identifier to obtain the corresponding video reconstruction auxiliary parameters.

[0165] It should be noted that, in this embodiment, the data in the semantic communication configuration library can be set with an expiration date or expiration time. Optionally, the data expiration date can be dynamically adjusted based on the session's activity level, and session configurations that have not been used for a long time can be automatically cleaned up to free up storage space. Simultaneously, the semantic communication configuration library supports persistent data storage and backup, ensuring rapid recovery after system failures or restarts.

[0166] In this embodiment, the generation and allocation of the unified semantic identifier can be completed at the initial stage of session establishment. For example, when the first terminal establishes a session with the second terminal, the IMS network element can generate a unique unified semantic identifier and store it in association with the video reconstruction auxiliary parameters of this session. During subsequent communication, this unified semantic identifier will be continuously transmitted in SIP signaling, enabling the network side to quickly obtain and use the same video reconstruction auxiliary parameters based on this identifier.

[0167] It should be noted that, in this embodiment, to address the timeliness issues of configuration changes and group adjustments, the semantic communication configuration library can support a configuration update mechanism. Optionally, when a user modifies video reconstruction auxiliary parameters or adjusts contact groups, a unified semantic identifier can be regenerated upon the next session establishment, or the parameters associated with the existing identifier can be updated. Simultaneously, the semantic communication configuration library can record configuration version information and update timestamps to ensure the use of the latest configuration parameters.

[0168] Step S402: The video reconstruction auxiliary parameters associated with the unified semantic identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal.

[0169] In this embodiment, after the IMS network element obtains the video reconstruction auxiliary parameters associated with the unified semantic identifier from the semantic communication configuration library, it determines the parameters as the video reconstruction auxiliary parameters corresponding to the first terminal and uses them in the subsequent video call data generation process. This avoids the process of temporarily generating video reconstruction auxiliary parameters and significantly improves the efficiency of video reconstruction.

[0170] This embodiment achieves efficient management and reuse of video reconstruction auxiliary parameters by introducing a unified semantic identifier mechanism. The unified semantic identifier serves as a unique identifier for a set of video reconstruction auxiliary parameter configurations, enabling the network side to quickly locate and obtain the required parameter configurations during a session, avoiding the overhead of repeated queries and ad-hoc generation. This parameter management mechanism based on unified semantic identifiers not only simplifies the session establishment process but also provides a technical foundation for subsequent cross-modal session configuration inheritance and mode switching within a session, significantly improving the overall performance and user experience of the semantic communication system.

[0171] In one feasible implementation, the preset extended field in the header of the first SIP invitation request carries a unified semantic identifier.

[0172] In this embodiment, the preset extended field can be a custom field in the SIP message header, such as the "X-Semantic-ID" field. This field is used to transmit a unified semantic identifier during SIP signaling interactions.

[0173] This embodiment achieves compatibility extension with the existing SIP protocol by carrying a unified semantic identifier in a preset extended field of the SIP message header. This extension method does not require modification of the core structure of the SIP protocol; it only requires adding a custom field to the message header, thus exhibiting good compatibility and scalability.

[0174] In one feasible implementation, the above-described communication method may further include steps S403-S405: Step S403: If the first SIP invitation request does not carry a unified semantic identifier, obtain the group identifier of the first terminal for the second terminal from the pre-configured customer group database, and obtain the video reconstruction auxiliary parameters associated with the group identifier from the video reconstruction configuration database based on the group identifier.

[0175] In this embodiment, the group identifier refers to the group identifier set by the user of the first terminal for the user of the second terminal. The user of the first terminal can manage multiple contacts in groups, such as strangers, colleagues, classmates, relatives, etc. Each group can correspond to different privacy levels and video reconstruction strategies.

[0176] In this embodiment, the customer group database is used to store the contact group information of the first terminal user. The video reconstruction configuration database is used to store the video reconstruction configuration information corresponding to each group.

[0177] It should be noted that, in this embodiment, the video reconstruction auxiliary parameters associated with the group identifier can be configured in two ways: The first configuration method involves configuring complete, ready-to-use video reconstruction auxiliary parameters directly for each group. In this method, the video reconstruction configuration database stores complete parameter configurations, including all necessary information such as the digital human avatar, scene background, and lighting parameters. Once the group identifier is obtained, the complete parameter configuration can be read directly from the database without additional processing.

[0178] The second configuration method involves setting only the configuration level or label for the group, rather than the complete auxiliary parameters. In this case, it is necessary to first obtain the user's complete video reconstruction auxiliary parameters, and then, based on the configuration level or label corresponding to the group identifier, trim or adjust the complete parameters to obtain truly usable video reconstruction auxiliary parameters. For example, for the "Strangers Group," only the basic digital human image might be retained, removing personalized features, or even converting it to a cartoon style; for the "Relatives Group," a digital human image that is closer to the user's real appearance might be used.

[0179] In this embodiment, the customer group database and the video reconstruction configuration database can support dynamic updates and real-time synchronization. Optionally, when a user modifies the contact group or video reconstruction configuration on the terminal, the updated data can be uploaded to the database through a synchronization mechanism to ensure that the network side can obtain the latest configuration information.

[0180] Step S404: The video reconstruction auxiliary parameters associated with the group identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal; a unique unified semantic identifier is generated, and the unified semantic identifier is associated with the video reconstruction auxiliary parameters corresponding to the first terminal and stored in the semantic communication configuration library.

[0181] In this embodiment, after obtaining the video reconstruction auxiliary parameters associated with the group identifier, the IMS network element identifies these parameters as the video reconstruction auxiliary parameters corresponding to the first terminal. Then, it generates a unique unified semantic identifier and associates this identifier with the video reconstruction auxiliary parameters, storing it in the semantic communication configuration library.

[0182] In this embodiment, when storing the unified semantic identifier and video reconstruction auxiliary parameters in the semantic communication configuration library, the validity period or expiration time of the storage can be set. Optionally, the storage validity period can be set according to the expected duration of the session, such as 24 hours, 7 days, or 30 days; or the validity period can be dynamically adjusted according to the user's activity status, extending the validity period for active sessions and shortening the validity period for inactive sessions.

[0183] Step S405: Insert the unified semantic identifier into the message header of the first SIP invitation response. The first SIP invitation response carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the first terminal and the second terminal during IMS communication both carry the unified semantic identifier. The types of SIP requests include SIP invitation requests and SIP message requests, and the types of SIP responses include SIP invitation responses and SIP message responses.

[0184] In this embodiment, when the IMS network element generates the first SIP invitation response, it inserts a Uniform Semantic Identifier (USI) into a preset extended field in the response message header. After receiving the response, the first terminal parses and saves the USI, and carries the USI in the message headers of all SIP requests and SIP responses sent in subsequent IMS communications with the second terminal.

[0185] In this embodiment, the end-to-end transmission mechanism of the unified semantic identifier can ensure the consistency and coherence of video reconstruction parameters throughout the session. Optionally, the first terminal can store the unified semantic identifier in a local cache and automatically add the identifier each time SIP signaling is sent; alternatively, it can maintain the unified semantic identifier in the session context to ensure continuous use throughout the session lifecycle.

[0186] It is important to note that in this embodiment, to address the timeliness issues of configuration changes and group adjustments, multiple mechanisms can be employed to ensure the use of the latest configuration parameters. Optionally, version information or timestamps can be embedded in the unified semantic identifier. When an expired version or an outdated timestamp is detected, a reconfiguration process is triggered. Alternatively, the user's latest group configuration can be checked at each session establishment. If a group change is detected, the unified semantic identifier and video reconstruction auxiliary parameters are regenerated. Furthermore, a configuration change notification mechanism can be set up to send an update notification to the relevant session when the user modifies the group or configuration, triggering parameter re-acquisition and identifier updates.

[0187] In this embodiment, the validity period management of the unified semantic identifier can adopt multiple strategies. Optionally, a fixed validity period can be set, such as 7 days or 30 days, which needs to be regenerated after expiration; the validity period can also be dynamically adjusted according to the activity level of the session, extending the validity period for contacts who communicate frequently and shortening the validity period for contacts who have not communicated for a long time; the validity period can also be set according to the frequency of configuration changes, shortening the validity period for users whose configurations change frequently and extending the validity period for users whose configurations are stable.

[0188] This embodiment achieves intelligent configuration and personalized customization of video reconstruction auxiliary parameters by obtaining them based on the customer group database and video reconstruction configuration database, even when the first SIP invitation request does not carry a unified semantic identifier (i.e., at the initial stage of session establishment). The group identifier mechanism allows users to configure different video reconstruction parameters for different groups based on the relationships and privacy requirements of different contacts, meeting diverse user needs for personalized experience and privacy protection. By supporting both full parameter configuration and configuration level / tag methods, this embodiment provides a flexible configuration strategy, allowing direct use of preset full parameters or dynamic adjustment based on user-defined basic parameters to adapt to different business scenarios and user preferences. Generating a unified semantic identifier and associating it with the video reconstruction auxiliary parameters lays the foundation for rapid parameter retrieval in subsequent sessions, avoiding repeated querying and configuration processes, and significantly improving system efficiency. Inserting the unified semantic identifier into the message header of the SIP invitation response and ensuring that all subsequent IMS communication SIP signaling carries this identifier achieves end-to-end transmission and consistent application of video reconstruction auxiliary parameters, ensuring the consistency and continuity of video reconstruction parameters throughout the session. To address the timeliness issues of configuration changes and group adjustments, this embodiment employs multiple mechanisms, including version management, dynamic checks, and change notifications, to ensure the use of the latest configuration parameters and avoid user experience problems caused by outdated configurations. This parameter configuration and management mechanism, based on group identifiers and unified semantic identifiers, not only enables intelligent configuration and efficient management of video reconstruction parameters but also provides a technical foundation for cross-modal session configuration inheritance and mode switching within a session, significantly improving the overall performance and user experience of the semantic communication system.

[0189] In one feasible implementation, the above-described communication method may further include steps S10 to S40: Step S10: Receive a SIP message request from the first terminal, wherein the SIP message request is a request from the first terminal to initiate a message call to the second terminal.

[0190] In this embodiment, a SIP message request refers to a request from the first terminal to send an instant message to the second terminal via the IMS network. The SIP message request uses the MESSAGE method. The request line contains the SIP URI of the second terminal as the request target. The message header contains standard SIP header fields that identify the message sender, receiver, session unique identifier, request sequence number, and routing path. The message body contains the content of the instant message, which can be text, images, emoticons, or other data.

[0191] In this embodiment, after the user triggers the sending of an instant message operation, the first terminal generates a SIP message request and sends it to the IMS core network through the IMS access network.

[0192] Step S20: In response to the SIP message request, it is determined that the first terminal has enabled the semantic service function, and the SIP message request does not carry a unified semantic identifier. The first terminal obtains the group identifier for the second terminal from the pre-configured customer group database, and based on the group identifier, the video reconstruction auxiliary parameters associated with the group identifier are obtained from the video reconstruction configuration database.

[0193] The content of this step is similar to step S403 in the previous embodiment, and can be referred to the previous description. This embodiment will not repeat it here.

[0194] Step S30: The video reconstruction auxiliary parameters associated with the group identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal; a unique unified semantic identifier is generated, and the video reconstruction auxiliary parameters corresponding to the first terminal are associated with the unified semantic identifier and stored in the semantic communication configuration library.

[0195] The content of this step is similar to step S404 in the previous embodiment, and can be referred to the previous description. This embodiment will not repeat it here.

[0196] Step S40: Insert the unified semantic identifier into the message header of the SIP message request, and send the SIP message request carrying the unified semantic identifier to the second terminal. The SIP message request carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the second terminal in IMS communication with the first terminal both carry the unified semantic identifier.

[0197] In this embodiment, the IMS network element inserts a unified semantic identifier into the header of the SIP message request before forwarding the SIP message request to the second terminal.

[0198] In this embodiment, after a SIP message request carrying a Uniform Semantic Identifier (USI) is sent to the second terminal, the second terminal parses and saves the USI upon receiving the message request. In subsequent IMS communications with the first terminal, the second terminal includes this USI in the headers of both the SIP request and SIP response messages.

[0199] It should be noted that, in this embodiment, the end-to-end transmission mechanism of the unified semantic identifier can ensure the consistency and coherence of video reconstruction parameters in subsequent sessions. Optionally, the second terminal can store the unified semantic identifier in a local cache and automatically add the identifier each time SIP signaling is sent; alternatively, it can maintain the unified semantic identifier in the session context to ensure continuous use throughout the session lifecycle.

[0200] This embodiment introduces a unified semantic identifier mechanism into the SIP message request processing flow, enabling configuration inheritance and seamless switching from 5G messaging sessions to 5G new video call sessions. When a user initiates instant messaging communication via a SIP message request, the IMS network element obtains video reconstruction auxiliary parameters based on the customer packet database and video reconstruction configuration database, and generates a unified semantic identifier for associated storage. This mechanism ensures that the required video reconstruction auxiliary parameters are pre-configured when the user subsequently switches to a video call. Inserting the unified semantic identifier into the message header of the SIP message request and ensuring that all subsequent SIP signaling in IMS communication carries this identifier achieves end-to-end transmission and consistent application of video reconstruction auxiliary parameters. When a user switches from a 5G messaging session to a 5G new video call session, they can directly inherit the video reconstruction auxiliary parameters configured in the previous session without reconfiguration, significantly simplifying the business process and improving the user experience. This unified semantic identifier pre-configuration mechanism based on SIP message requests not only enables cross-modal session configuration inheritance, but also provides a technical foundation for mode switching in a session. This allows users to initiate video calls at any time during message communication, and the system can immediately use the pre-configured parameters, significantly shortening the video call startup time and significantly improving the overall performance and user experience of the semantic communication system.

[0201] In one feasible implementation, the above communication method may further include step S50: Step S50: Receive the SIP message response returned by the second terminal based on the SIP message request carrying a unified semantic identifier, and send the SIP message response to the first terminal.

[0202] Among them, the SIP message response carries a unified semantic identifier, which ensures that the SIP requests and SIP responses sent by the first terminal and the second terminal in IMS communication both carry a unified semantic identifier.

[0203] In this embodiment, after receiving a SIP message request carrying a Uniform Semantic Identifier (USI), the second terminal generates a SIP message response and returns it to the IMS network element. The SIP message response is typically a 200 OK response, indicating that the message has been successfully received and processed. The header of the SIP message response also carries a USI to ensure the continuous transmission of the identifier in bidirectional communication.

[0204] In this embodiment, before forwarding the SIP message response to the first terminal, the IMS network element can perform necessary processing on the response content. Optionally, it can check whether the unified semantic identifier is correctly carried in the response; if it is not carried or the identifier is incorrect, it can be supplemented or corrected. Alternatively, other content in the response can be filtered or transformed to meet specific service requirements.

[0205] In this embodiment, the IMS network element sends the processed SIP message response to the first terminal. After receiving the response, the first terminal parses and saves the unified semantic identifier to ensure that the identifier is continuously used in subsequent IMS communications with the second terminal.

[0206] This embodiment achieves efficient management and reuse of video reconstruction auxiliary parameters by introducing a unified semantic identifier mechanism. The unified semantic identifier serves as a unique identifier for a set of video reconstruction auxiliary parameter configurations, remaining unchanged across different session types, making cross-modal session configuration inheritance possible. When a user switches from a 5G messaging session to a new 5G video call session, they can directly inherit the video reconstruction auxiliary parameters configured in the previous session without reconfiguration, significantly simplifying the business process. Simultaneously, the unified semantic identifier supports mode switching within the same session, pre-configuring video reconstruction auxiliary parameters at the beginning of the session. This allows for immediate use when switching from a voice call to a video call by turning on the camera at any time, eliminating the need for temporary parameter generation and significantly shortening the startup time of the video semantic call.

[0207] Please refer to Figure 4 This application proposes a communication method according to a fourth embodiment.

[0208] In the fourth embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0209] In this embodiment, the communication method described above may further include step S500: In step S500, in response to the first SIP invitation request, it is determined that the first terminal has enabled the semantic service function but does not have the video semantic extraction capability. A second SIP invitation response is sent to the first terminal. The second SIP invitation response carries a second SDP response, and the second SDP response carries fourth information. The fourth information is used to instruct the first terminal to send only audio data from the media data collected in the video call.

[0210] In this embodiment, unlike the first and second embodiments, this embodiment addresses a scenario where the first terminal lacks video semantic extraction capabilities. When the IMS network element determines that the first terminal has enabled semantic service functions but lacks video semantic extraction capabilities, it sends a second SIP invitation response to the first terminal. The second SIP invitation response carries a second SDP response containing fourth information.

[0211] It is important to note that the fourth information differs fundamentally in function from the first information in the first embodiment. The first information instructs the first terminal to convert the video data collected during the video call into video semantic features; that is, the terminal is responsible for video semantic extraction and then sends the video semantic features and audio data together to the IMS network element. The fourth information, however, explicitly instructs the first terminal to send only audio data from the collected media data, without sending any video data at all. This difference reflects the differentiated adaptation strategy of this embodiment for terminals with different capabilities: for terminals with video semantic extraction capabilities, the solution of the first embodiment is adopted, where the terminal is responsible for extracting and sending video semantic features; for terminals without video semantic extraction capabilities, the solution of this embodiment is adopted, where the terminal is only responsible for audio collection and transmission, and no longer performs video semantic extraction.

[0212] This embodiment achieves differentiated service support for terminals with different capabilities by using fourth information to instruct the terminal to send only audio data from the collected media data when the first terminal lacks video semantic extraction capabilities. Compared with the first and second embodiments, the core innovation of this embodiment lies in the adaptation to terminal capabilities: when the terminal lacks video semantic extraction capabilities, audio data can be sent by adjusting the media transmission strategy. The driving force of the digital human image no longer relies on video semantic features, but is based on audio data analysis and preset parameters. This design not only ensures the universality of semantic communication services for terminals with different capabilities, but also achieves the rational allocation of computing resources, enabling terminals without semantic extraction capabilities to enjoy the privacy protection and personalized experience brought by semantic communication services, significantly expanding the scope of application and user group of semantic communication services.

[0213] Step S600: Receive audio data sent by the first terminal.

[0214] Step S700: Obtain the video reconstruction auxiliary parameters corresponding to the first terminal, wherein the video reconstruction auxiliary parameters include the digital human image.

[0215] This step is the same as step S410 in the previous embodiment, and can be referred to the previous description. This embodiment will not repeat it here.

[0216] In step S800, based on the video reconstruction auxiliary parameters and the audio data sent by the first terminal, second video call data is generated and sent to the second terminal.

[0217] Among them, the active display of digital human figures in the second video call data includes one of the following: At least one of the following elements of a digital human avatar—its posture, actions, and facial expressions—is preset and remains unchanged; At least one of the digital human's posture, actions, and expressions changes based on preset change conditions; The digital human's posture, movements, and expressions are driven by emotional information, sentiment information, and / or lip-sync data obtained from the analysis of audio data.

[0218] In this embodiment, when at least one of the digital human's posture, actions, and expressions is preset and unchanged, the IMS network element uses fixed parameters to drive the corresponding part of the digital human when generating the second video call data. This driving method is suitable for scenarios where dynamic performance requirements are not high, or scenarios where users want to maintain a specific image, and has the advantages of simple implementation and low computational load.

[0219] When at least one of the digital human's posture, actions, or expressions changes based on preset conditions, the IMS network element dynamically adjusts the corresponding part of the digital human's image according to these preset conditions (such as time period, scene conditions, user preferences, etc.) when generating the second video call data. This driving method is suitable for scenarios that require a certain degree of dynamic performance but do not require real-time driving, increasing the naturalness and vividness of the digital human's image through periodic or conditional changes.

[0220] When at least one of the digital human's posture, actions, and expressions is driven by emotional information, affective information, and / or lip-sync data obtained from audio data analysis, the IMS network element first analyzes the audio data to extract emotional information, affective information, and / or lip-sync data when generating the second video call data, and then drives the corresponding part of the digital human based on this information.

[0221] It should be noted that, compared with the first and second embodiments, this embodiment differs in that it lacks video semantic features. In the first and second embodiments, the driving force of the digital human image mainly relies on video semantic features, including structured information such as the person's posture, actions, and expressions; however, in this embodiment, since the terminal does not have the ability to extract video semantics and does not send video data, the network side cannot obtain video semantic features. Therefore, the driving force of the digital human image relies entirely on audio data analysis and preset parameters.

[0222] At the lip-sync data extraction level, IMS network elements can use voice-driven lip-sync technology to analyze audio data, extract the speaker's lip-sync information, and achieve basic lip-sound synchronization.

[0223] At the level of emotion information extraction, IMS network elements can use voice emotion recognition technology to analyze audio data, extract the speaker's emotional state, and use it to drive the facial expressions of the digital human avatar.

[0224] At the level of emotional information extraction, IMS network elements can use voice emotion analysis technology to perform deeper analysis of audio data, extract the speaker's emotional tendencies and intensity, and use this to drive the overall performance of the digital human avatar.

[0225] In this embodiment, when the first terminal lacks video semantic extraction capabilities, the fourth information instructs the terminal to send only audio data from the collected media data. Based on the audio data and video reconstruction auxiliary parameters, second video call data is generated, achieving differentiated service support for terminals with different capabilities. Compared to the first and second embodiments, the core innovation of this embodiment lies in the adaptation to terminal capabilities: when the terminal lacks video semantic extraction capabilities, the media transmission strategy is adjusted so that the driving force of the digital human image no longer relies on video semantic features, but is based on audio data analysis or preset parameters. This differentiated adaptation mechanism based on terminal capabilities expands the applicability of semantic communication services and significantly improves the overall performance and user experience of the semantic communication system.

[0226] In one feasible implementation, the second SDP response further includes fifth information, which instructs the first terminal to send the audio data it collected during the video call to the target address.

[0227] The above-mentioned step S600, receiving audio data sent by the first terminal, may include step S610: Step S610: Receive audio data sent by the first terminal based on the target address.

[0228] This embodiment is similar to step S310 in the previous embodiment, and can be referred to the previous description. This embodiment will not be repeated here.

[0229] In one feasible implementation, the above communication method may further include step S510: In step S510, in response to the first SIP invitation request, it is determined that the first terminal has enabled semantic service function but does not have video semantic extraction capability, and a third SIP invitation request is sent to the second terminal. The third SIP invitation request carries a second SDP proposal.

[0230] The second SDP proposal includes a sixth message, which instructs the second terminal to receive second video call data from the target address during a video call.

[0231] This embodiment is similar to the embodiment corresponding to step S210 in the previous embodiments, and can be referred to the previous description, which will not be repeated here.

[0232] In one feasible implementation, when the first terminal has enabled the semantic service function and the second terminal has not enabled the semantic service function, the second SDP response also includes third video media line direction information for indicating that only video media lines are sent, third audio media line direction information for indicating bidirectional transmission of audio media lines, and third data media line direction information for indicating that data media lines are inactive.

[0233] The first SDP proposal also includes fourth video media line direction information for indicating bidirectional transmission of video media lines, fourth audio media line direction information for indicating bidirectional transmission of audio media lines, and fourth data media line direction information for indicating inactive data media lines.

[0234] This implementation method is similar to the implementation methods in the foregoing embodiments, and can be referred to the previous description, so it will not be repeated here.

[0235] In one feasible implementation, before obtaining the video reconstruction auxiliary parameters corresponding to the first terminal in step S700 above, steps S701 to S702 may be included: Step S701: If the first SIP invitation request carries a unified semantic identifier, obtain the video reconstruction auxiliary parameters associated with the unified semantic identifier from the preset semantic communication configuration library.

[0236] Step S702: The video reconstruction auxiliary parameters associated with the unified semantic identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal.

[0237] This implementation method is the same as the implementation method corresponding to steps S401 to S402 in the previous embodiments, and can be referred to the previous description, so it will not be repeated here.

[0238] In one feasible implementation, the preset extended field in the header of the first SIP invitation request carries a unified semantic identifier.

[0239] This implementation method is the same as the implementation method in the foregoing embodiments, and can be referred to the previous description, so it will not be repeated here.

[0240] In one feasible implementation, the above-described communication method may further include steps S703-S705: Step S703: If the first SIP invitation request does not carry a unified semantic identifier, obtain the group identifier of the first terminal for the second terminal from the pre-configured customer group database, and obtain the video reconstruction auxiliary parameters associated with the group identifier from the video reconstruction configuration database based on the group identifier.

[0241] Step S704: The video reconstruction auxiliary parameters associated with the group identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal; a unique unified semantic identifier is generated, and the unified semantic identifier is associated with the video reconstruction auxiliary parameters corresponding to the first terminal and stored in the semantic communication configuration library.

[0242] Step S705: Insert the unified semantic identifier into the message header of the second SIP invitation response. The second SIP invitation response carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the first terminal and the second terminal during IMS communication carry the unified semantic identifier. The types of SIP requests include SIP invitation requests and SIP message requests, and the types of SIP responses include SIP invitation responses and SIP message responses.

[0243] This embodiment is similar to the embodiment corresponding to steps S403 to S405 in the previous embodiments, and can be referred to the previous description, which will not be repeated here.

[0244] To further clarify the complete deployment architecture and typical application scenarios of the technical solution of this application in operator networks, this application provides a specific embodiment, and combines it with... Figure 5 , Figure 6 and Figure 7 The communication method and the internal logical architecture of the IMS network element in the above embodiments of this application are described in detail.

[0245] In this specific embodiment, the IMS network element is a network entity that simultaneously includes control layer logic functions and media layer logic functions. The terminal layer interacts with the IMS network element through standard SIP / SDP signaling, realizing the end-to-end network deployment of generative communication.

[0246] like Figure 5 As shown, this specific embodiment is based on a standard IMS network and adopts a three-layer architecture of control layer—media layer—terminal layer: 1. Terminal layer, including the sending terminal UE-A (i.e., the first terminal) and the receiving terminal UE-B (i.e., the second terminal).

[0247] The terminal layer uses the native client as a unified entry point, and the client contains at least the following functional modules: Terminal capability discovery module: This module detects whether a terminal possesses a Neural Processing Unit (NPU) and rendering engine capabilities, and generates a terminal capability identifier (i.e., capability indication information). This identifier is carried in standard SIP signaling and reported to the IMS network element.

[0248] Multimodal feature extraction module (NPU): The terminal uses this module to convert the video data collected during video calls into video semantic features and upload them to the IMS network element.

[0249] 2. Control layer: The part of the IMS network element responsible for signaling interaction, deployed on the operator's application server (AS), internally containing the following functional sub-modules that work together: The Policy Generation and Signaling Control Submodule acts as a Back-to-Back User Agent (B2BUA) to process and terminate SIP signaling transactions, serving as the sole exit point for external signaling interaction at the control layer. It is responsible for inserting a unified semantic identifier into the semantic extension fields (i.e., preset extension fields) of the standard SIP signaling message header to carry the communication intent. Based on the communication policy output by the intent mapping submodule and the identity parameters trimmed by the identity semantic submodule, it performs semantic-level extension and asymmetric negotiation on the session description protocol, dynamically reconstructing the media topology—including but not limited to adjusting the direction of media description lines, migrating video generation anchor points to network-side edge nodes, or restoring standard bidirectional transmission based on terminal capabilities. It maintains a unified semantic identifier and session state mapping table (i.e., a semantic communication configuration library), associating discrete message sessions (e.g., 5G message sessions) with real-time call sessions (e.g., 5G new call sessions), enabling video reconstruction auxiliary parameters such as user classification policies, enterprise image parameters, and privacy configurations in preceding messages to be automatically inherited into subsequent real-time calls, achieving zero-configuration cross-modal migration.

[0250] The intent mapping submodule, serving as a policy decision support module within the control layer, is responsible for receiving query requests from the policy generation and signaling control submodules. It parses the end-user's communication intent and matches it with corresponding user-level policies, scenario tags, privacy control levels, and interface interaction templates. This can be implemented through a pre-configured rule base, enterprise policy interface queries, or an intelligent decision engine, without being limited to a specific implementation form. In other words, it retrieves the group identifier of the first terminal for the second terminal from a pre-configured customer group database and determines the communication policy based on this group identifier.

[0251] Identity Semantics Submodule: As the identity semantic support module within the control layer, it is responsible for generating identity semantic vectors (Identity Embedding) based on the image data authorized by the enterprise or end user, and performing dimensional clipping according to the privacy control level issued by the policy generation and signaling control submodule. In strict mode, only structural features are output, while in normal mode, complete features are output, that is, the final digital human image is output.

[0252] Collaboration Relationship: After the policy generation and signaling control submodule terminates the SIP transaction, it calls the intent mapping submodule as needed to obtain the communication policy and the identity semantic submodule to obtain the trimmed identity parameters. Finally, it generates a complete communication control policy (i.e., video reconstruction auxiliary parameters) and injects it into the signaling layer. The above calls are collaborations between modules within the control layer, and externally, it is still presented as a single B2BUA entity.

[0253] 3. Media Layer: This is the part of the IMS network element responsible for media generation. It is deployed on the media processing nodes of the operator's network side and can be flexibly configured on edge cloud nodes or central cloud resource pools according to service latency, computing power requirements, and deployment conditions. Internally, it includes the following functional sub-modules that work together: The synchronization and disaster recovery submodule is responsible for media layer data alignment, quality monitoring, and reconstruction strategy control. Its specific functions are as follows: 1) Double-buffered lock-and-synchronize: The system receives video semantic features transmitted via the data channel and audio data transmitted via the audio media channel from the receiving terminal. Semantic feature buffer queues and audio buffer queues are maintained separately, and inter-stream alignment is performed based on a unified timestamp benchmark. The aligned composite data unit serves as input to the cloud-based logical reconstruction engine. In other words, it receives video semantic features sent by the first terminal and audio data collected by the first terminal during a video call; timestamp alignment is performed based on the video semantic features and audio data.

[0254] 2) Network Quality Assessment: Real-time monitoring of transmission latency, packet loss rate, and jitter to generate a network quality rating. In other words, it obtains the current network quality parameters for the video call, including transmission latency and / or packet loss rate.

[0255] 3) Dynamic Degradation and Reconstruction Strategy Control: When network quality deteriorates, a speed-down instruction (reducing the semantic feature update frequency) is sent back to the sending terminal; simultaneously, a reconstruction strategy update instruction is sent to the cloud-based logical reconstruction engine, including but not limited to switching to a lightweight model, reducing output resolution, or triggering preset image interpolation, to ensure reconstruction continuity. That is, based on network quality parameters, video rendering strategy parameters are determined, including video resolution and / or the parameter magnitude of the generation model that generates the first video call data; when the network quality parameters indicate that the current network quality of the video call is a first network quality, a first frequency instruction is sent to the first terminal, instructing the first terminal to update the video semantic features at the first frequency; when the network quality parameters indicate that the current network quality of the video call is a second network quality, a second frequency instruction is sent to the first terminal, instructing the first terminal to update the video semantic features at the second frequency; wherein, the first network quality is lower than the second network quality, and the first frequency is lower than the second frequency.

[0256] The cloud-based logic reconstruction engine is responsible for receiving aligned composite data units from the synchronization and disaster recovery submodules, and simultaneously receiving trimmed identity semantic vectors from the control layer identity semantic submodule as generation conditions to ensure that the output image is consistent with the communication subject. It also supports the following reconstruction modes: 1) Real-time streaming reconstruction (with NPU path): Based on aligned structured semantic features and audio Mel-spectral features, combined with injected identity semantic vectors, frame-level generative reconstruction is performed, and the output video stream is pushed to the receiving terminal. That is, video reconstruction auxiliary parameters corresponding to the first terminal are obtained, including the digital human image; based on the video reconstruction auxiliary parameters, as well as the timestamp-aligned video semantic features and audio data, the first video call data is generated.

[0257] 2) Pre-set Image-Driven Reconstruction (No NPU Path): This method only receives the audio media stream and a pre-set identity semantic vector. Based on audio-viewpoint mapping, it drives the pre-set image and generates a standard video stream, which is then pushed to the receiving terminal. Specifically, it obtains the video reconstruction auxiliary parameters corresponding to the first terminal, including the digital human image; and generates second video call data based on the video reconstruction auxiliary parameters and the audio data sent by the first terminal.

[0258] like Figure 6 As shown in this specific embodiment, the specific business process of 5G new call real-time semantic communication is as follows: Step 1, Standard Signaling Access and Service Triggering: The UE-A terminal initiates a standard SIP INVITE request, and the SDP proposes to declare audio and video media capabilities normally. At the same time, the UE-A terminal capability discovery module detects terminal capability identifiers such as local NPU and rendering engine version (X-Terminal-Capability) and sends them back to the control layer via the SIP INVITE request for network-side policy decision-making.

[0259] All SIP requests are routed to the control layer via the IMS core network. The policy generation and signaling control submodules in the control layer query user subscription data. If semantic communication service is confirmed to be enabled: extract the terminal capability identifier from INVITE and proceed to step 2 to generate semantic policy; If not activated: Transmit according to the standard audio and video call process, and terminate the subsequent processing of this specific embodiment.

[0260] Step 2, Semantic Policy Generation: The policy generation and signaling control submodule acts as a back-to-back user agent to terminate SIP transactions for UE-A and extract the terminal capability identifier from INVITE; it calls the intent mapping submodule to obtain the scene label, privacy control level and preset image of UE-A for UE-B; and it calls the identity semantic submodule to perform dimensional pruning of the identity semantic vector according to the privacy control level—only structural features are output in strict mode, and complete features are output in normal mode.

[0261] The policy generation and signaling control submodule inserts a unified semantic identifier carrying the communication intent into the semantic extension field of the SIP message header and generates a communication control policy.

[0262] Step 3, SDP asymmetric negotiation and media topology reconstruction: The policy generation and signaling control submodule regenerates the SDP Answer for UE-A: confirming that the audio and data channels are bidirectional, the video channel is only transmitted, and requiring UE-A to stop transmitting video bitstreams and switch to uplink video semantic features.

[0263] Simultaneously, the policy generation and signaling control submodule regenerates the SDP Offer for UE-B: the network address points to the output address of the cloud-based logical reconstruction engine (i.e., the target address); the video and audio channels are bidirectional, and the data channel is set to inactive. UE-B returns an SDP Answer, confirming receipt of the reconstructed video from the network side.

[0264] Thus, the media topology has been reconstructed from the traditional point-to-point dual transmission and reception to a star-shaped topology: UE-A uplink video semantic features + audio data, reconstructed in the cloud and then sent down to UE-B.

[0265] Step 4, Real-time Video Semantic Feature Extraction and Uplink: The UE-A's multimodal feature extraction module calls the terminal NPU to perform inference on the facial images captured by the camera, outputting skeletal keypoints, expression blendshape weights, and head pose Euler angles, packaging them into structured video semantic feature frames. These video semantic feature frames are streamed to the network via the DataChannel, with the original face pixels remaining on the terminal. Audio data is transmitted synchronously via a standard RTP channel.

[0266] Step 5, Double Buffer Alignment, Network Monitoring and Dynamic Rendering Strategy Generation: The media layer synchronization and disaster recovery submodule receives the uplink data from UE-A, maintains the semantic feature buffer queue and the audio buffer queue, performs double queue phase-locked synchronization according to the timestamp identifier of the unified clock reference, and atomically prepares to send it to the cloud logic reconstruction engine after successful matching.

[0267] The synchronization and disaster recovery submodule monitors network transmission quality (end-to-end latency, packet loss rate, jitter) in real time, and combines this with terminal capability parameters (screen resolution, decoding performance, frame rate support limit) issued by the control layer and the rendering material requirements corresponding to scene tags to comprehensively generate dynamic rendering strategy instructions, i.e., video rendering strategy parameters. In other words, it obtains the current network quality parameters of the video call, including transmission latency and / or packet loss rate; and determines the video rendering strategy parameters based on these network quality parameters.

[0268] The rendering strategy instructions include at least: Video resolution and frame rate; The selection of the generation model (standard model / lightweight model / ultra-lightweight model) is also known as the parameter scale of the generation model that generates the first video call data. Whether to enable preset image interpolation to ensure continuity; Injection strength of identity semantic vector.

[0269] The synchronization and disaster recovery submodule sends the aligned semantic features, audio Mel spectrum features, and dynamic rendering strategy instructions to the cloud-based logic reconstruction engine.

[0270] The cloud-based logic reconstruction engine performs frame-level generative reconstruction based on the above inputs. It injects the cropped identity semantic vector issued by the control layer's identity semantic submodule as a generation condition to ensure identity consistency. It strictly follows the output parameters (resolution, frame rate, model specifications) of the rendering strategy instructions to encode a standard video stream, which is then pushed to the UE-B via the RTP channel. In other words, based on the video rendering strategy parameters, video reconstruction auxiliary parameters, and timestamp-aligned video semantic features and audio data, it generates the first video call data and sends it to the second terminal.

[0271] Step 6, Adjusting the Network Adaptive Strategy: The synchronization and disaster recovery submodule continuously monitors network quality changes: when transmission latency, packet loss rate, or jitter exceeds a preset threshold, it dynamically adjusts rendering strategy instructions (reduces resolution, switches to a lightweight model, and enables preset image interpolation); simultaneously, it sends a speed-down feedback to the UE-A (reduces the semantic feature update frequency). That is, when the network quality parameters indicate that the current network quality of the video call is a first network quality, it sends a first frequency instruction to the first terminal, which instructs the first terminal to update the video semantic features at a first frequency; when the network quality parameters indicate that the current network quality of the video call is a second network quality, it sends a second frequency instruction to the first terminal, which instructs the first terminal to update the video semantic features at a second frequency; wherein, the first network quality is lower than the second network quality, and the first frequency is lower than the second frequency.

[0272] When the network quality recovers to a good level, the dynamic callback rendering strategy instruction is switched to high-fidelity mode.

[0273] The cloud-based logic refactoring engine responds to policy instructions in real time, adjusts the output video quality, and enables network-adaptive generative media transmission.

[0274] like Figure 7 As shown in this specific embodiment, the specific business process of 5G message asynchronous semantic communication is as follows: Scenario Definition: Enterprise A (such as a bank, e-commerce platform, or government agency) sends asynchronous message notifications (such as bill reminders, logistics updates, or service notifications) to User B. The enterprise backend only issues a simplified business identifier and basic content variables, without any interface configuration, interaction logic, or media parameters. The policy generation and signaling control submodule acts as the operator-side policy hub, terminating SIP transactions and calling internal functional modules. It injects semantically extended fields and differentiated rendering policy packages into standard SIP messages, enabling the same service type to present completely different interactive interfaces and subsequent communication paths for different users. After completing the interaction in the message interface, User B can seamlessly trigger a real-time call by clicking a button. The user classification, enterprise policy, and business context in the preceding messages are automatically inherited to the real-time call session through a unified semantic identifier, achieving zero-configuration cross-modal migration.

[0275] Step 1, Minimalist Message Initiation and Standard Routing: The enterprise backend calls the Chatbot platform API (Application Programming Interface) to initiate a SIP message request, which only carries the following: Recipient identifier (user B's communication number); Business type identifier (e.g., bill reminder, logistics notification, government service guidance); Basic content variables (such as bill amount, payment date, tracking number, and item name).

[0276] The enterprise side does not include any interface layout definitions, button configurations, card styles, digital human avatar parameters, or interaction flow scripts.

[0277] The above information is assembled into a standard SIP message by the Chatbot platform, enters the IMS core network through the MaaP gateway, and is routed to the control layer.

[0278] Step 2, Control Layer Policy Aggregation and Semantic Injection: Step 2.1, POLICY calls the MAP (map table) to query enterprise user policies: The Policy Generation and Signalling Control submodule (POLICY) extracts the sender's enterprise identifier, service type, and B's number, and then calls the Intent Mapping submodule (MAP). MAP stores pre-defined user hierarchy policy templates for the enterprise, such as interface template identifiers and interaction permission configurations for VIP (Very Important Person) users, regular users, new customers, and high-risk users. MAP returns the hierarchy tags and scenario matching results for B.

[0279] Step 2.2, POLICY calls the ID (identifier) ​​to perform corporate image trimming: The strategy generation and signaling control submodule calls the identity semantics submodule (ID) to extract the Identity Embedding of A from the enterprise identity semantics library, and performs cropping according to the image precision strategy returned by MAP: strict: Only outputs general structure parameters; normal: Outputs the complete identity semantic vector.

[0280] The cropped corporate image parameters are retained by POLICY in the session state for use in subsequent cross-modal real-time calls.

[0281] Step 2.3, POLICY calls SYNC (Synchronization) to execute discrete encapsulation: The strategy generation and signaling control submodule encapsulates the basic content variables and differentiated rendering strategy instruction packages (interface layout, button events, interactive scripts) into a single semantic message unit by the synchronization and disaster recovery module (SYNC), and marks it with a unified semantic identifier.

[0282] Step 2.4, POLICY injection of semantic extension fields: The policy generation and signaling control submodule, acting as a back-to-back user agent, terminates enterprise-side SIP transactions and inserts a unified semantic identifier into the SIP MESSAGE header field for the enterprise. The message body encapsulates semantic message units of the SYNC output.

[0283] Step 3, Differentiated distribution and terminal capability feedback: The policy generation and signaling control submodule routes the SIP message to UE-B via the IMS core network. UE-B's terminal capability discovery module detects the screen size and rendering engine version, generates a terminal adaptation identifier, and sends it back to POLICY via SIP 200 OK and IMS.

[0284] Step 4, Terminal Differentiated Rendering: The native UE-B client parses semantic extension fields, identifies unified semantic identifiers, determines user classification, and loads corresponding templates: VIP: Brand avatar + rich media card + "Video Consultant / Delay / Confirmation"; Standard: Standard avatar + message bubble + "Confirm / Question"; New customers: Carousel card + "Claim benefits / Apply now"; Risk: Warning sign + red highlight + single confirmation; Step 5, Cross-modal semantic inheritance (associated through a unified semantic identifier): When a user clicks "Video Advisor" on the message interface, the UE-B initiates a SIP INVITE request and carries the unified semantic identifier of the preceding message.

[0285] The IMS core network routes the INVITE to the control layer. The policy generation and signaling control submodule queries the session state mapping table, matches the previous message state, and extracts the retained information. User classification tags in the intent mapping submodule; Corporate image parameters retained after trimming the identity semantic submodule; Business context (billing amount, repayment status, etc.); Enterprise's subsequent communication strategy.

[0286] The policy generation and signaling control submodule sends the retained corporate image parameters to the cloud-based logic reconstruction engine (RENDER) to regenerate the SDP. The RENDER then performs real-time streaming reconstruction and outputs the video stream to UE-B. All semantic configurations in the preceding messages are automatically carried over, eliminating the need to reconfigure user levels, privacy policies, or corporate image parameters.

[0287] It should be noted that the above embodiments / implementations are only used to assist in understanding this application and do not constitute a limitation on the communication method of this application. Any simple modifications based on this technical concept are all within the protection scope of this application.

[0288] In addition, please refer to Figure 8 , Figure 8 This is a schematic diagram of the IMS network element structure of the hardware operating environment involved in the communication method in the embodiments of this application.

[0289] This application also provides an IMS network element, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the steps of the communication method in the above embodiments.

[0290] The following is for reference. Figure 8 It shows a schematic diagram of the structure of an IMS network element suitable for implementing the embodiments of this application. Figure 8The IMS network element shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0291] like Figure 8 As shown, an IMS network element may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the IMS network element. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays, speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tape, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the IMS network element to exchange data via wireless or wired communication with other devices. Although the diagram shows IMS network elements with various systems, it should be understood that it is not required to implement or have all of the systems shown. Alternatively, more or fewer systems may be implemented.

[0292] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0293] The IMS network element provided in this application, employing the communication method described in the above embodiments, can solve the technical problem of how to implement generative communication in IMS. Compared with related technologies, the beneficial effects of the IMS network element provided in this application are the same as those of the communication method provided in the above embodiments, and other technical features of this IMS network element are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0294] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0295] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above claims.

[0296] In addition, this application also provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the steps of the communication method in the above embodiments.

[0297] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency, etc., or any suitable combination thereof.

[0298] The aforementioned computer-readable storage medium may be included in the IMS network element or may exist independently without being assembled into the IMS network element.

[0299] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by an IMS network element, the IMS network element causes the IMS network element to: receive a first SIP invitation request from a first terminal, wherein the first SIP invitation request is a request from the first terminal to initiate a video call to a second terminal; in response to the first SIP invitation request, determine that the first terminal has enabled semantic service functions and has video semantic extraction capabilities, and send a first SIP invitation response to the first terminal, wherein the first SIP invitation response carries a first SDP response, wherein the first SDP response includes first information, the first information being used to instruct the first terminal to convert the video data collected during the video call into video semantic features; receive the video semantic features sent by the first terminal, and the audio data collected by the first terminal during the video call; perform timestamp alignment based on the video semantic features and audio data, and generate first video call data based on the timestamp-aligned video semantic features and audio data, and send it to the second terminal.

[0300] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or connected to an external computer, such as via the Internet.

[0301] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0302] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0303] The computer-readable storage medium provided in this application stores computer-readable program instructions for performing the steps of the communication method in the above embodiments, which can solve the technical problem of how to implement generative communication in IMS. Compared with related technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the communication method provided in the above embodiments, and will not be repeated here.

[0304] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the communication method described in the above embodiments.

[0305] The computer program product provided in this application solves the technical problem of how to implement generative communication in IMS. Compared with related technologies, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the communication methods provided in the above embodiments, and will not be repeated here.

[0306] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A communication method, comprising: Receive a first SIP invitation request from a first terminal, wherein the first SIP invitation request is a request from the first terminal to initiate a video call to the second terminal; In response to the first SIP invitation request, it is determined that the first terminal has enabled semantic service function and has video semantic extraction capability. A first SIP invitation response is sent to the first terminal. The first SIP invitation response carries a first SDP response. The first SDP response includes first information. The first information is used to instruct the first terminal to convert the video data collected in the video call into video semantic features. Receive the video semantic features sent by the first terminal, and the audio data collected by the first terminal during the video call; Based on the video semantic features and the audio data, timestamp alignment is performed, and based on the timestamp-aligned video semantic features and audio data, first video call data is generated and sent to the second terminal.

2. The method as described in claim 1, characterized in that, The first SDP response also includes second information, which instructs the first terminal to send the video semantic features and the audio data collected by the first terminal during the video call to the target address. The step of receiving the video semantic features sent by the first terminal and the audio data collected by the first terminal during the video call includes: Based on the target address, the video semantic features sent by the first terminal and the audio data collected by the first terminal during the video call are received.

3. The method as described in claim 2, characterized in that, The method further includes: In response to the first SIP invitation request confirming that the first terminal has enabled semantic service functions and has video semantic extraction capabilities, a second SIP invitation request is also sent to the second terminal, the second SIP invitation request carrying the first SDP proposal; The first SDP proposal includes third information, which is used to instruct the second terminal to receive the first video call data from the target address during the video call.

4. The method as described in claim 3, characterized in that, The first SDP response further includes first video media line direction information for indicating that only video media lines are transmitted, first audio media line direction information for indicating that audio media lines are transmitted bidirectionally, and first data media line direction information for indicating that only data media lines are received. The first SDP proposal also includes second video media line direction information for indicating bidirectional transmission of video media lines, second audio media line direction information for indicating bidirectional transmission of audio media lines, and second data media line direction information for indicating inactive data media lines.

5. The method as described in claim 1, characterized in that, The video semantic features are structural semantic features that characterize the activity information of people in the video data. The activity information of people includes at least one of the following: person's posture, person's actions, and person's facial expressions. The step of generating the first video call data based on the timestamp-aligned video semantic features and audio data includes: Obtain the video reconstruction auxiliary parameters corresponding to the first terminal, wherein the video reconstruction auxiliary parameters include a digital human image; Based on the video reconstruction auxiliary parameters, as well as the video semantic features and audio data after timestamp alignment, the first video call data is generated; The video semantic features are used to drive the digital human image to perform active displays, and the active displays include the display of at least one of the following: human posture, human actions, and human expressions.

6. The method as described in claim 5, characterized in that, Before the step of obtaining the video reconstruction auxiliary parameters corresponding to the first terminal, the method further includes: If the first SIP invitation request carries a unified semantic identifier, the video reconstruction auxiliary parameters associated with the unified semantic identifier are obtained from the preset semantic communication configuration library; The video reconstruction auxiliary parameters associated with the unified semantic identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal.

7. The method as described in claim 6, characterized in that, The preset extended field in the header of the first SIP invitation request carries the unified semantic identifier.

8. The method as described in claim 6, characterized in that, The method further includes: If the first SIP invitation request does not carry a unified semantic identifier, the group identifier of the first terminal for the second terminal is obtained from the pre-configured customer group database, and based on the group identifier, the video reconstruction auxiliary parameters associated with the group identifier are obtained from the video reconstruction configuration database. The video reconstruction auxiliary parameters associated with the group identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal; a unique unified semantic identifier is generated, and the unified semantic identifier is associated with the video reconstruction auxiliary parameters corresponding to the first terminal and stored in the semantic communication configuration library; The unified semantic identifier is inserted into the message header of the first SIP invitation response. The first SIP invitation response carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the first terminal and the second terminal during IMS communication both carry the unified semantic identifier. The types of the SIP requests include SIP invitation requests and SIP message requests, and the types of the SIP responses include SIP invitation responses and SIP message responses.

9. The method as described in claim 6, characterized in that, The method further includes: Receive a SIP message request from a first terminal, wherein the SIP message request is a request from the first terminal to initiate a message call to a second terminal; In response to the SIP message request confirming that the first terminal has enabled semantic service function, and the SIP message request does not carry a unified semantic identifier, the group identifier of the first terminal for the second terminal is obtained from the pre-configured customer group database, and based on the group identifier, the video reconstruction auxiliary parameters associated with the group identifier are obtained from the video reconstruction configuration database. The video reconstruction auxiliary parameters associated with the group identifier are determined as the video reconstruction auxiliary parameters corresponding to the first terminal; a unique unified semantic identifier is generated, and the video reconstruction auxiliary parameters corresponding to the first terminal are associated with the unified semantic identifier and stored in the semantic communication configuration library; The unified semantic identifier is inserted into the message header of the SIP message request, and the SIP message request carrying the unified semantic identifier is sent to the second terminal. The SIP message request carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the second terminal and the first terminal during IMS communication both carry the unified semantic identifier.

10. The method as described in claim 9, characterized in that, The method further includes: Receive the SIP message response returned by the second terminal based on the SIP message request carrying the unified semantic identifier, and send the SIP message response to the first terminal; The SIP message response carries the unified semantic identifier, and the SIP message response carrying the unified semantic identifier ensures that the SIP requests and SIP responses sent by the first terminal and the second terminal during IMS communication both carry the unified semantic identifier.

11. The method as described in claim 7, characterized in that, The step of generating the first video call data based on the video reconstruction auxiliary parameters, as well as the timestamp-aligned video semantic features and audio data, includes: Obtain the current network quality parameters of the video call, wherein the network quality parameters include transmission latency and / or packet loss rate; Based on the network quality parameters, video rendering strategy parameters are determined, wherein the video rendering strategy parameters include the video resolution and / or the parameter magnitude of the generation model that generates the first video call data, and the higher the parameter magnitude, the greater the computing power resource requirement of the generation model. Based on the video rendering strategy parameters, the video reconstruction auxiliary parameters, and the video semantic features and audio data after timestamp alignment, the first video call data is generated.

12. The method as described in claim 11, characterized in that, After the step of obtaining the current network quality parameters of the video call, the method further includes: When the network quality parameter indicates that the current network quality of the video call is a first network quality, a first frequency instruction is sent to the first terminal. The first frequency instruction is used to instruct the first terminal to update the video semantic features at a first frequency. When the network quality parameter indicates that the current network quality of the video call is the second network quality, a second frequency instruction is sent to the first terminal, the second frequency instruction being used to instruct the first terminal to update the video semantic features at the second frequency; Wherein, the first network quality is lower than the second network quality, and the first frequency is lower than the second frequency.

13. The method as described in claim 1, characterized in that, The method further includes: In response to the first SIP invitation request, which determines that the first terminal has enabled semantic service function but does not have video semantic extraction capability, a second SIP invitation response is sent to the first terminal. The second SIP invitation response carries a second SDP response, and the second SDP response carries fourth information. The fourth information is used to instruct the first terminal to send only audio data from the media data collected in the video call. Receive audio data sent by the first terminal; Obtain the video reconstruction auxiliary parameters corresponding to the first terminal, wherein the video reconstruction auxiliary parameters include a digital human image; Based on the video reconstruction auxiliary parameters and the audio data sent by the first terminal, second video call data is generated and sent to the second terminal, wherein the active display of the digital human image in the second video call data includes one of the following: At least one of the following: the pose, actions, and expressions of the digital human figure is preset and unchanging; At least one of the digital human's posture, actions, and expressions changes based on preset change conditions. The digital human's posture, actions, and expressions are driven by at least one of the emotional information, sentiment information, and / or lip-sync data obtained from the analysis of the audio data.

14. An IMS network element, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the communication method as described in any one of claims 1 to 13.

15. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the communication method as described in any one of claims 1 to 13.

16. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the communication method as described in any one of claims 1 to 13.