Call setup method, communication method, server and terminal

CN122845748APending Publication Date: 2026-09-29MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610953852.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,上述辅助手段虽然可以一定程度上改善远程沟通效果,但均存在明显的局限性:手语翻译人员协助存在高昂的人力成本,且远程通话前需要提前协调手语翻译人员的工作时间,无法实现及时远程通话;短信文字交流与远程通话相比,输入速度慢、接收方回复不及时,交流效率低;手语识别软件通常仅能进行简单的手语识别,难以实现实时的双向转换和流畅通信

Benefits of technology

[0015]在本申请的技术方案中,服务器接收第一终端发送的第一视频流;识别第一视频流中的人物行为特征,并基于识别结果生成第一音频流;基于运营商通话,发送第一音频流至第二终端;和/或,服务器基于运营商通话,接收第二终端发送的第二音频流;识别第二音频流中的语音内容,并基于识别结果生成第二视频流,第二视频流包含用于表征语音内容的人物行为特征;发送第二视频流至第一终端。如此,本申请实施例由运营商的服务器集中提供通话中肢体语言与语音之间的转换功能,听障用户与健听用户无需进行翻译准备即可建立无障碍的运营商通话,提高了听障用户的远程沟通效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845748A_ABST
    Figure CN122845748A_ABST
Patent Text Reader

Abstract

The application provides a call establishment method, a communication method, a server and a terminal. The call establishment method comprises the following steps: receiving a first video stream sent by a first terminal; identifying a human behavior feature in the first video stream, and generating a first audio stream based on the identification result; sending the first audio stream to a second terminal based on an operator call; and / or, the server receives a second audio stream sent by the second terminal based on the operator call; identifying speech content in the second audio stream, and generating a second video stream based on the identification result, wherein the second video stream contains a human behavior feature for representing the speech content; and sending the second video stream to the first terminal. The application centrally provides a conversion function between body language and speech in a call by a server of an operator, so that a hearing-impaired user and a normal-hearing user can establish an operator call without barrier without translation preparation, and the remote communication efficiency of the hearing-impaired user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to communication technology, and more particularly to a call establishment method, a communication method, a server, and a terminal. Background Technology

[0002] Remote communication typically relies on the transmission and reception of sound. Hearing users can make voice calls using landlines, mobile phones, and other communication terminals. However, for hearing-impaired users with hearing difficulties, this method of remote communication cannot meet their communication needs. To address the barriers to remote communication between hearing-impaired and hearing users, relevant technologies typically employ methods such as sign language interpreters, text messaging, and sign language recognition software to assist in remote communication between the two groups.

[0003] However, while the aforementioned auxiliary methods can improve remote communication to some extent, they all have significant limitations: the assistance of sign language interpreters incurs high labor costs, and their work schedules need to be coordinated in advance before remote calls, making timely remote communication impossible; text messaging is slower to input and receives delayed responses, resulting in lower communication efficiency compared to remote calls; sign language recognition software typically only performs simple sign language recognition and struggles to achieve real-time two-way conversion and smooth communication. Therefore, enabling hearing-impaired users to establish real-time, barrier-free two-way remote calls with hearing users is a pressing technical problem that needs to be solved in this field. Summary of the Invention

[0004] In view of this, embodiments of this application provide a call establishment method, a communication method, a server, and a terminal.

[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a call establishment method, applied to a server, the method comprising: Receive the first video stream sent by the first terminal; Identify the behavioral characteristics of people in the first video stream, and generate a first audio stream based on the identification results; Based on the carrier call, the first audio stream is sent to the second terminal; And / or, Based on the operator's call, receive the second audio stream sent by the second terminal; The speech content in the second audio stream is identified, and a second video stream is generated based on the identification result. The second video stream contains human behavioral features used to characterize the speech content. Send the second video stream to the first terminal.

[0006] In the above scheme, the behavioral characteristics of the person include sign language gestures.

[0007] In the above scheme, after recognizing the speech content in the second audio stream, the method further includes: Based on the recognition results, corresponding text information is generated; the text information is then sent to the first terminal.

[0008] The method in the above scheme further includes: Receive the first request sent by the first terminal; Based on the first request, establish a packet-switched network connection with the first terminal; The first video stream and / or the second video stream are transmitted based on an established packet-switched network connection.

[0009] The method in the above scheme further includes: Receive a second request sent by the first terminal, wherein the second request carries the identifier of the second terminal; Based on the second request, a carrier call is established with the second terminal via the Session Initialization Protocol (SIP).

[0010] Secondly, embodiments of this application provide a communication method applied to a first terminal, the method comprising: In the accessibility call mode, a first interface is displayed, which includes one or more icons representing the terminal to be called; In response to the user's selection of a target icon, a request is made to establish an accessibility call with the second terminal corresponding to the target icon; After the barrier-free call is established, the second interface is displayed; An image is rendered in a first area of ​​the second interface. The image displayed in the first area contains user-identifiable human behavioral characteristics, which characterize the communication content of the user of the second terminal.

[0011] The method in the above scheme further includes: Capture video facing the user; Based on the video facing the user, an image is rendered in the second area of ​​the second interface.

[0012] The method in the above scheme further includes: Text information is displayed in the third area of ​​the second interface, and the text information is used to represent the communication content of the user of the second terminal.

[0013] Thirdly, embodiments of this application provide a server, the server comprising: A processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a first terminal, the first terminal comprising: A processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method as described in the second aspect.

[0015] In the technical solution of this application, the server receives a first video stream sent by a first terminal; identifies the behavioral characteristics of people in the first video stream and generates a first audio stream based on the identification result; sends the first audio stream to a second terminal based on the operator's call; and / or, the server receives a second audio stream sent by the second terminal based on the operator's call; identifies the speech content in the second audio stream and generates a second video stream based on the identification result, the second video stream containing behavioral characteristics of people used to characterize the speech content; and sends the second video stream to the first terminal. Thus, in this embodiment of the application, the operator's server centrally provides the conversion function between body language and speech during a call, allowing hearing-impaired users and hearing users to establish barrier-free operator calls without the need for translation preparation, thereby improving the efficiency of remote communication for hearing-impaired users. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the first process of the call establishment method according to an embodiment of this application; Figure 2 This is a schematic diagram of the second process of the call establishment method according to an embodiment of this application; Figure 3 This is a schematic diagram of the call flow between a hearing-impaired user and a hearing user according to an embodiment of this application; Figure 4 This is a flowchart illustrating the communication method according to an embodiment of this application; Figure 5 This is a schematic diagram of the first interface of an embodiment of this application; Figure 6 This is a schematic diagram of the second interface in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of the first terminal and the server; Figure 8 This is a schematic diagram of the interaction process between the first terminal, the server, and the second terminal. Figure 9 A schematic diagram of the device for establishing a call; Figure 10 This is a schematic diagram of the communication device. Figure 11 This is a schematic diagram of the server structure; Figure 12 This is a schematic diagram of the structure of the first terminal. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] With the development of communication technology, users can communicate in real time through various forms of remote communication, such as carrier calls, internet voice calls, and video calls, greatly improving communication efficiency and convenience. Remote calls usually rely on the acquisition and transmission of the other party's voice. However, for hearing-impaired users (also known as deaf-mute users), they cannot hear sounds or express themselves verbally. Therefore, conventional voice calls cannot meet the remote communication needs of hearing-impaired users.

[0019] Hearing-impaired users typically communicate through sign language and other body language, while most hearing users do not possess sign language skills or cannot accurately understand the meaning of sign language representations. Therefore, there are obstacles to remote communication between hearing-impaired and hearing users. To address these obstacles, related technologies employ methods such as sign language interpreters, text messaging, sign language recognition software, and specialized equipment to assist hearing-impaired users in making remote calls with hearing users.

[0020] However, while the aforementioned assistive methods can improve the effectiveness of remote communication for hearing-impaired users to some extent, they all have significant limitations. For instance, video calls combined with sign language interpreters can translate the conversation between hearing-impaired and hearing users into understandable speech or sign language. However, sign language interpreter assistance incurs high labor costs, and coordinating interpreters' schedules before remote calls is necessary, hindering timely communication. Furthermore, the quality of interpreters' translations is directly affected by their professional competence and fatigue levels. Compared to remote calls, text-based communication methods such as SMS and instant messaging suffer from slow input speeds and difficulty in comprehension, failing to meet communication needs for rapid and accurate communication, resulting in low efficiency. Additionally, text message notifications are typically triggered only once on the communication terminal, making it difficult for the recipient to notice notifications promptly, leading to poor communication timeliness. While sign language recognition software can recognize simple sign language, it cannot translate the speech of hearing users into speech for hearing-impaired users, making real-time, two-way, fluent communication difficult. Dedicated communication devices for hearing-impaired users are expensive and uncomfortable to wear, hindering their widespread adoption within the hearing-impaired community.

[0021] Therefore, in response to the aforementioned urgent problem of difficulties in remote communication between hearing-impaired users and hearing users, this application provides a call establishment method and a communication method, aiming to improve the efficiency of remote communication for hearing-impaired users.

[0022] For example, the call establishment method provided in this application embodiment is applied to a server, including a first call establishment method and / or a second call establishment method.

[0023] The method for establishing the first call is as follows: Figure 1 As shown, the method includes: Step 101: Receive the first video stream sent by the first terminal.

[0024] Step 102: Identify the behavioral characteristics of the people in the first video stream and generate the first audio stream based on the identification results.

[0025] Step 103: Based on the operator call, send the first audio stream to the second terminal.

[0026] Here, the server in this embodiment is used to establish a call between the first terminal and the second terminal. The server may take the form of a cloud server, a network server, an application server, etc., and this embodiment does not specifically limit it.

[0027] In some embodiments, the server is deployed in a telecommunications operator's network or connected to the telecommunications operator's IP Multimedia Subsystem (IMS) network. The IMS network is the core network used by the telecommunications operator to support multimedia services such as voice calls.

[0028] Here, the first terminal and the second terminal in this application embodiment are communication terminals, specifically the communication parties that establish a call, that is, the call users use the first terminal and the second terminal to conduct remote calls respectively.

[0029] In this embodiment, the first terminal is a communication terminal used by a hearing-impaired user, also known as an accessible communication terminal. The form of the first terminal includes, but is not limited to, devices that support real-time video communication, such as mobile phones, tablets, and desktop computers. This embodiment does not specifically limit the form of the first terminal. In this embodiment, the user on the first terminal side is a hearing-impaired user.

[0030] In this embodiment, the second terminal is a regular communication terminal, that is, a communication terminal that supports real-time voice communication and does not need to support additional call functions. It can be understood as a regular communication terminal that supports basic dialing and answering functions. The form of the second terminal includes, but is not limited to, devices that support real-time voice communication such as landline phones, mobile phones, tablets, and desktop computers. This embodiment does not specifically limit this. In this embodiment, the user on the second terminal side is a hearing user.

[0031] It should be noted that the call established in this application embodiment is a remote call between a hearing-impaired user and a hearing user, which is different from a regular remote call between hearing users. The call established in this application embodiment will be referred to as an accessibility call in the following description. For example, the method further includes: in response to an accessibility call request sent by the first terminal, establishing an accessibility call between the first terminal and the second terminal.

[0032] It is understandable that the first terminal is the call initiator of the accessibility call, and the server determines that the users in this call include hearing-impaired users based on the received accessibility call request.

[0033] Here, a carrier call refers to a real-time communication connection established through the core network of a telecommunications operator. Carrier calls include, but are not limited to, VoLTE (Voice over Long-Term Evolution) calls and VoNR (Voice over New Radio) calls based on the IMS network, circuit-switched 2G / 3G voice calls, and PSTN (Public Switched Telephone Network) fixed-line calls. Unlike OTT (Over-The-Top) calls such as VoIP (Voice over Internet Protocol) based on the public internet, carrier calls are characterized by session control by the telecommunications operator's core network, and the receiving end does not need to install a corresponding communication application to answer the call.

[0034] It should be noted that in this embodiment of the application, the server and the second terminal communicate via a carrier call.

[0035] It is understandable that when the server establishes an accessibility call, it establishes a carrier call with the second terminal.

[0036] Here, the first video stream sent by the first terminal is generated from the real-time user video captured by the first terminal. The first terminal is equipped with a video capture device facing the user, such as a front-facing camera, which captures real-time user video and encodes the captured real-time user video into the first video stream.

[0037] It should be noted that, since hearing-impaired users cannot communicate by voice, in order to achieve the purpose of communication, during barrier-free calls, the hearing-impaired user, as the party in the call, expresses the content of the communication through body movements. The body movements of the hearing-impaired user are captured by the video capture device of the first terminal facing the user and generate real-time user video. The real-time user video is then encoded and sent to the server in the form of a first video stream.

[0038] It is understandable that the first video stream contains the behavioral characteristics of the user on the first terminal side, including the user's body language used to express communication content.

[0039] It should be noted that, since the user on the second terminal is a hearing user, they cannot recognize and understand the communication content expressed by the body movements of the user on the first terminal. Therefore, in the accessibility call, after the server receives the first video stream, it does not directly forward the first video stream to the second terminal. Instead, it decodes the first video stream, identifies the behavioral characteristics of the people in the decoded first video stream, and converts the identified behavioral characteristics into speech content that the hearing user can understand.

[0040] It should be noted that, in order for the server to accurately convert the identified human behavior features into speech content, the server presets the correspondence between human behavior features and speech content or text information; in some embodiments, human behavior features include sign language gestures.

[0041] Here, sign language gestures have standardized semantic correspondences, meaning that each sign language gesture corresponds to a specific word or semantic meaning. Thus, the server can accurately convert the recognized sign language gestures into corresponding speech or text content based on the preset correspondences.

[0042] It is understandable that sign language is a common form of communication that most hearing-impaired users have mastered. In this embodiment of the application, hearing-impaired users communicate remotely with the other side of the hearing-free user through sign language gestures.

[0043] In some embodiments, identifying human behavior features in a first video stream includes: identifying human behavior features in the decoded first video stream according to a first model, and generating a recognition result. Correspondingly, generating a first audio stream based on the recognition result includes: converting the recognition result into speech content according to a second model, and generating a first audio stream based on the speech content.

[0044] In some embodiments, the recognition results of a person's behavioral characteristics are text information.

[0045] Here, the first model is an artificial intelligence (AI) behavior recognition model, including but not limited to: deep learning-based convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM), and Transformers.

[0046] In one example, the first model can employ a two-stage processing architecture: the first stage is spatial feature extraction, where the server inputs the video frames obtained after decoding the first video stream into a CNN-based spatial feature extractor to extract spatial features of key user parts and poses within the video frames; the second stage is temporal feature modeling, where the server inputs the spatial feature sequences of consecutive video frames into an LSTM-based temporal model to capture the temporal dynamic changes of spatial features and identify corresponding gesture sequences and other forms of human behavior features. After identifying the human behavior features, the server maps the recognition results to corresponding text information.

[0047] It should be noted that the server pre-determines the correspondence between human behavior features and language content or text information. The first model can be trained based on this correspondence, so that the first model can convert the identified human behavior features into corresponding text information.

[0048] Here, the second model is a text-to-speech conversion model, including but not limited to end-to-end text-to-speech (TTS) models.

[0049] For example, the second model adopts an encoder-decoder architecture, combined with an attention mechanism, which can generate Mel spectra from the text information of the recognition results frame by frame, and then synthesize the Mel spectra into speech waveforms through a vocoder.

[0050] It should be noted that the processing method of the server converting the first video stream into the first audio stream is not limited to the processing method based on the first model and the second model described above, and the embodiments of this application do not specifically limit it in this regard.

[0051] It is understood that in this embodiment, the server on the core network side provides accessible calling services. When the first terminal communicates with the second terminal based on the established accessible call, no translation assistance such as sign language interpreters is required. Specifically, the hearing-impaired user on the first terminal side expresses the communication content through sign language gestures and uploads the corresponding real-time user video to the server. The server performs real-time translation and conversion of body language and speech based on its built-in data processing function, and sends the first audio stream generated after conversion to the second terminal, so that the user on the second terminal side can accurately understand the communication content expressed by the hearing-impaired user on the first terminal side through speech.

[0052] It should be noted that the user on the second terminal, as the recipient of the accessible call, perceives no difference between the accessible call and a regular call. The audio stream received by the second terminal is completely identical to the voice content in a regular carrier call. Users do not need to learn new operations or adapt to a special interface; they can participate in the accessible call in the same way as answering a regular phone call. Throughout the call, the user is completely unaware of the background body language-voice conversion processing.

[0053] It should be noted that in this embodiment, the server and the second terminal communicate based on operator calls, enabling barrier-free calls to access the mature operator core network. The second terminal used by the hearing user does not need any modification or configuration of a dedicated application to access barrier-free calls. Moreover, during barrier-free calls, the hearing user still receives the same voice information as in regular calls. The conversion between body language and voice is imperceptible throughout the call and is not limited by terminal functions, enabling hearing-impaired users to conduct real-time and convenient remote communication with users of any ordinary communication terminal.

[0054] It is understood that in this embodiment of the application, the operator's server centrally provides the function of converting body language and speech during a call, so that hearing-impaired users and hearing users can establish barrier-free operator calls without translation preparation, thereby improving the efficiency of remote communication for hearing-impaired users.

[0055] In some embodiments, establishing an unobstructed call between a first terminal and a second terminal includes: receiving a first request sent by the first terminal; and establishing a packet-switched network connection with the first terminal based on the first request.

[0056] Here, the aforementioned accessibility call request includes the first request.

[0057] Here, the first request is used to request the establishment of a packet-switched network connection with the server.

[0058] Here, after the server establishes a packet-switched network connection with the first terminal, a media channel is formed for transmitting video streams.

[0059] The first video stream is transmitted based on an established packet-switched network connection.

[0060] It should be noted that during the barrier-free call, real-time video stream transmission occurs between the first terminal and the server. Although the operator's call supports video call services, the first terminal establishing a barrier-free call based on the operator's call would not only change the original establishment mechanism of the operator's call, but also have the problem of high cost of operator video call services. Therefore, in the embodiment of this application, the first terminal establishes a packet-switched network connection with the server during the establishment of the barrier-free call, so as to serve as the media channel for video stream transmission during the subsequent call, thereby achieving low-cost and highly compatible real-time video stream transmission.

[0061] In some embodiments, establishing a packet-switched network connection with the first terminal includes: establishing a Web Real-time Communications (WebRTC) connection with the first terminal.

[0062] Here, WebRTC is a technology standard that enables real-time audio and video communication in web browsers. Since most mainstream web browsers natively support WebRTC, terminals with web browsers installed do not need additional video communication software; they can directly conduct real-time video communication through the browser.

[0063] It is understood that the first terminal in this application embodiment can install a dedicated accessibility call application to realize the corresponding accessibility call function, or it can initiate an accessibility call through an installed web browser without installing a dedicated application. That is, the first terminal can be a communication terminal with a dedicated application and / or a web browser installed, without the need for hardware modification of the first terminal.

[0064] In some embodiments, the first request includes a proposed session description protocol offer (Offer SDP).

[0065] Here, the Session Description Protocol (SDP) is a protocol format used to describe multimedia session parameters, such as media type, encoding format, and network address. During the WebRTC connection establishment process, the two communicating parties need to exchange SDP information through a signaling channel to negotiate media configuration. The SDP information sent by the initiator is called the Offer SDP, and the SDP information replied by the receiver is called the Answer SDP. Through this offer / answer mechanism, the two communicating parties can reach an agreement on media parameters, thereby establishing a WebRTC media channel.

[0066] Understandably, after receiving the Offer SDP sent by the first terminal, the server generates the corresponding AnswerSDP and then establishes a WebRTC connection.

[0067] In some embodiments, establishing an unobstructed call between a first terminal and a second terminal further includes: receiving a second request sent by the first terminal, the second request carrying the identifier of the second terminal; and establishing a carrier call with the second terminal via SIP based on the second request.

[0068] Here, the aforementioned accessibility call request includes the second request.

[0069] The identifier for the second terminal is used to indicate the second terminal to which the call is to be established. This identifier can be a telephone number or other communication address that can uniquely identify the second terminal.

[0070] Here, SIP is an application-layer signaling control protocol used to establish, modify, and terminate multimedia sessions. Based on the SIP interaction process, the server establishes a carrier call with the second terminal corresponding to the second request through the carrier's core network, such as the IMS network, thereby realizing the server's call establishment and session management for the second terminal.

[0071] It is understood that during the barrier-free call establishment process in this application embodiment, the server establishes a packet-switched network connection with the first terminal and an operator call with the second terminal based on the first request and the second request sent sequentially by the first terminal. The complete barrier-free call channel consists of the packet-switched network media channel and the operator call. Based on the signaling conversion and fusion conversion mechanism of "packet-switched network + SIP", the server provides body language-speech conversion services during the call, which reduces the real-time remote communication cost between hearing-impaired users and hearing users and improves the remote communication efficiency of hearing-impaired users.

[0072] The method for establishing the second call is as follows: Figure 2 As shown, the method includes: Step 201: Based on the operator's call, receive the second audio stream sent by the second terminal.

[0073] Step 202: Identify the speech content in the second audio stream and generate a second video stream based on the identification result. The second video stream contains human behavioral features used to characterize the speech content.

[0074] Step 203: Send the second video stream to the first terminal.

[0075] It is understandable that the specific connotations of the server, the first terminal, the second terminal, the operator's call, and the behavioral characteristics of the person in the second call establishment method are consistent with the similar descriptions in the first call establishment method, and will not be repeated here.

[0076] In this embodiment, the server and the second terminal communicate via a carrier call; in some embodiments, the server establishes a carrier call with the second terminal based on an accessibility call request sent by the first terminal.

[0077] Here, the second audio stream is generated based on real-time user speech collected on the second terminal side; during barrier-free calls, the user on the second terminal side dictates the communication content to the second terminal, and the sound acquisition device such as the microphone configured on the second terminal collects the real-time user speech, which is then encoded and sent to the server in the form of a second audio stream.

[0078] It is understandable that the second audio stream contains the voice information of the user on the second terminal side.

[0079] It should be noted that, since hearing-impaired users cannot communicate by voice, in order to achieve the purpose of communication, after receiving the second audio stream, the server does not directly forward the second audio stream to the first terminal. Instead, it decodes the second audio stream and identifies the speech content in the decoded second audio stream, and converts the identified speech content into human behavioral characteristics that the hearing-impaired user can understand.

[0080] Here, the character behavior characteristics are images of the virtual character's postures and actions, which are used to display to hearing-impaired users in video format on the first terminal side.

[0081] It should be noted that, in order for the server to accurately convert the recognized speech content into human behavior features, the server presets the correspondence between human behavior features and speech content or text information; in some embodiments, human behavior features include sign language gestures.

[0082] In some embodiments, identifying speech content in a second audio stream includes: identifying speech content in the decoded second audio stream according to a third model, and generating a recognition result. Correspondingly, generating a second video stream based on the recognition result includes: converting the recognition result into human behavior features according to a fourth model, and generating a second video stream based on the human behavior features.

[0083] In some embodiments, the recognition result of language content is text information.

[0084] In some embodiments, after recognizing the speech content in the second audio stream, the method further includes: generating corresponding text information based on the recognition result; and sending the text information to the first terminal.

[0085] Understandably, in order to help hearing-impaired users accurately understand the communication content expressed by hearing users, the text information corresponding to the voice content will also be sent to the first terminal based on the accessibility call.

[0086] Here, the third model is the speech-to-text conversion model, also known as the Automatic Speech Recognition (ASR) model or the Speech to Text (STT) model, including but not limited to: deep learning-based CNN, RNN, LSTM and Transformer.

[0087] Here, the fourth model is an AI behavior feature generation model, used to convert text information into a second video stream containing human behavior features. The fourth model includes, but is not limited to: a GPT-based autoregressive generation model, a diffusion-based video generation model, and a Transformer-based sequence-to-sequence model.

[0088] In one example, the fourth model can adopt a two-stage processing architecture: the first stage is text-to-action mapping, where the server inputs the text information of the recognition results into a Transformer-based text encoder to generate corresponding action token sequences (such as gesture sequences, body posture sequences, facial expression parameters, etc.); the second stage is video frame generation, where the server uses the action token sequences to generate a continuous second video stream containing human behavioral features through a renderer.

[0089] It should be noted that the server pre-determines the correspondence between human behavior features and language content or text information. The fourth model can be trained based on this correspondence, so that the fourth model can convert the recognized text information into the corresponding human behavior features.

[0090] Understandably, the behavioral characteristics generated by the fourth model correspond to standardized sign language gestures or body language expressions, enabling hearing-impaired users on the first terminal to understand the communication content expressed by users on the second terminal by watching the behavioral characteristics of the characters in the second video stream.

[0091] It should be noted that the server's processing method for converting the second audio stream into the second video stream is not limited to the processing methods based on the third and fourth models described above, and the embodiments of this application do not specifically limit this.

[0092] It is understood that in this embodiment, the server on the core network side provides accessible calling services. When the first terminal communicates with the second terminal based on the established accessible call, no translation assistance media such as sign language interpreters are required. Specifically, the user on the second terminal side normally dictates the communication content and uploads the corresponding real-time user voice to the server. The server performs real-time translation and conversion of speech and body language based on its built-in data processing function, and sends the generated second video stream to the first terminal, so that the hearing-impaired user on the first terminal side can accurately understand the communication content expressed by the user on the second terminal side through the behavioral characteristics of the person in the second video stream.

[0093] It should be noted that in this embodiment, the server and the second terminal communicate based on operator calls, enabling barrier-free calls to access the mature operator core network. The second terminal used by the hearing user does not need any modification or configuration of a dedicated application to access barrier-free calls. Moreover, during barrier-free calls, the hearing user still receives the same voice information as in regular calls. The conversion between body language and voice is imperceptible throughout the call and is not limited by terminal functions, enabling hearing-impaired users to conduct real-time and convenient remote communication with users of any ordinary communication terminal.

[0094] In some embodiments, sending a second video stream to a first terminal includes: sending the second video stream to the first terminal based on an established packet-switched network.

[0095] It is understandable that during the establishment of an accessible call, the first terminal establishes a packet-switched network connection with the server, and performs video stream interaction based on the established packet-switched network connection during the accessible call, thereby achieving low-cost and highly compatible video communication.

[0096] It is understood that in this embodiment of the application, the operator's server centrally provides the function of converting body language and speech during a call, so that hearing-impaired users and hearing users can establish barrier-free operator calls without translation preparation, thereby improving the efficiency of remote communication for hearing-impaired users.

[0097] In some embodiments, the server supports both the first call establishment method and the second call establishment method, that is, the server supports bidirectional conversion between body language and speech, which greatly reduces the difficulty of remote communication between hearing-impaired users and hearing users.

[0098] In one application example, the call flow between a hearing-impaired user and a hearing user is as follows: Figure 3 As shown, barrier-free communication between hearing-impaired users and hearing users is achieved based on the conversion of streaming data between the terminal and the server.

[0099] Based on the same technical concept, this application also provides a communication method, which is applied to a first terminal; such as Figure 4 As shown, the method includes: Step 401: In the accessibility call mode, a first interface is displayed, which includes one or more icons representing the terminal to be called.

[0100] Step 402: In response to the user's selection of the target icon, a request is made to establish an accessibility call with the second terminal corresponding to the target icon.

[0101] Step 403: After the barrier-free call is established, the second interface is displayed.

[0102] Step 404: Render an image in the first area of ​​the second interface. The image displayed in the first area contains user-identifiable human behavior characteristics, which represent the communication content of the user on the second terminal.

[0103] Here, the first terminal is a terminal device used by hearing-impaired users. The first terminal includes a display device in the form of a screen, which is used to display the corresponding interface based on user operations.

[0104] It is understood that the first terminal supports accessibility calling functionality. For example, the first terminal can install a dedicated accessibility calling application, or it can directly access accessibility calling mode through a web browser; this application embodiment does not specifically limit this.

[0105] Understandably, the first terminal responds to the user's operation of activating the accessibility call mode by displaying the first interface; wherein, the first interface is the call interface of the accessibility call, for example, the first interface can be the address book interface in the accessibility call mode.

[0106] Here, the user operation to activate the accessibility call mode can be the user operation of launching the application, the user operation of entering the corresponding service webpage, the user operation of completing the user login after launching the application, or the user operation of completing the user login after entering the corresponding service webpage.

[0107] Here, the icon representing the terminal to be called is used to identify the corresponding terminal. For example, this icon can be associated with information such as the terminal name, communication address, telephone number, or username of the terminal to be called; this embodiment of the application does not specifically limit this. The telephone number is an MSISDN (Mobile Station International ISDN Number).

[0108] In some embodiments, the terminal information to be called can be bound to the user's identity. After the user logs in in the barrier-free calling mode, the first interface displays the corresponding icon representing the terminal to be called.

[0109] In some embodiments, the terminal information to be called can be stored locally on the first terminal. In the barrier-free calling mode, the application directly calls the contact information stored on the first terminal and displays the corresponding icon representing the terminal to be called on the first interface.

[0110] It is understandable that the first terminal responds to the user's operation by displaying the first interface, and determines the target terminal for the request call based on the user's icon selection operation, which is the second terminal.

[0111] Here, the user on the second terminal side is a hearing user.

[0112] In some embodiments, requesting a second terminal corresponding to the target icon to establish an accessibility call specifically includes: requesting the server to establish a carrier call with the second terminal.

[0113] Here, the server is used to establish seamless communication between the first terminal and the second terminal; in some embodiments, the server is deployed in a telecommunications operator's network or connected to the telecommunications operator's IMS network.

[0114] Understandably, after the first terminal identifies the second terminal, it initiates an accessibility call request, asking the server to establish an accessibility call between the first and second terminals. This accessibility call relies on the core network of the telecommunications operator; that is, the server establishes an operator call with the second terminal through the operator's core network and interacts with the second terminal based on this established operator call.

[0115] Understandably, relying on the operator's core network, ordinary communication terminals that only support basic calling and answering functions (such as landlines, feature phones, etc.) can join the barrier-free call as a second terminal requesting a call, without any hardware modification or software configuration of the second terminal, nor the installation of any application.

[0116] In some embodiments, requesting the server to establish a carrier call with the second terminal includes: sending a second request to the server, the second request being used to request the server to establish a carrier call with the second terminal.

[0117] Here, the first terminal responds to the user's selection of the target icon, generates a corresponding second request, and sends the second request to the server.

[0118] Here, the second request carries the identifier of the second terminal.

[0119] In some embodiments, before sending the second request to the server, the method further includes: sending a first request to the server, the first request being used to request the first terminal to establish a packet-switched network connection with the server.

[0120] Here, the first terminal responds by displaying the first interface, generating the first request, and sending the first request to the server.

[0121] Understandably, the first request, as an accessibility call request, is sent after the hearing-impaired user on the first terminal side activates the accessibility call mode, requesting to establish a media channel with the server.

[0122] It is understandable that an unobstructed communication channel between the first terminal and the second terminal is established based on the packet-switched network connection established between the first terminal and the server and the operator call established between the server and the second terminal.

[0123] Here, the second interface is the call interface for accessible calls; it should be noted that, since the hearing-impaired user on the first terminal side cannot listen to the voice and recognize the communication content expressed by the voice, in this embodiment of the application, after the accessible call is established, the display interface of the first terminal switches to the second interface, and the hearing-impaired user obtains the communication content of the other user based on the image displayed in the first area of ​​the second interface.

[0124] Here, the image displayed in the first area contains user-recognizable human behavioral features, which represent body language; in some embodiments, the human behavioral features include sign language gestures.

[0125] In some embodiments, after an accessibility call is established, the method further includes: receiving a second video stream sent by a server; accordingly, rendering an image in a first area of ​​the second interface, including: rendering an image in the first area of ​​the second interface based on the second video stream.

[0126] Here, the first terminal receives the second video stream based on the established packet-switched network connection.

[0127] In some embodiments, the first terminal's request to establish an accessible call with the second terminal corresponding to the target icon is also used to instruct the server to provide a conversion function between body language and speech. Since the user on the second terminal side in the accessible call is a hearing user who usually does not have sign language skills, the accessible call supports the user on the second terminal side to make a call through speech. The server converts the second audio stream generated based on speech into a second video stream containing the characteristics of human behavior in real time, and then sends the second video stream to the first terminal so that the first terminal can display the communication content of the other user to the hearing-impaired user in the form of image rendering.

[0128] It is understandable that after the first terminal receives the second video stream, it decodes the second video stream and performs image rendering based on the decoded video frames.

[0129] It is understood that, in this embodiment of the application, after the barrier-free call is established, the first terminal renders an image containing human behavioral characteristics in real time in the first area of ​​the second interface. Hearing-impaired users can intuitively obtain the communication content of the other user by viewing the image, without relying on sign language interpreters or text input assistance, which reduces the communication threshold for hearing-impaired users and improves the convenience and real-time nature of remote communication.

[0130] In some embodiments, the method further includes: acquiring a video facing the user; and rendering an image in a second region of the second interface based on the video facing the user.

[0131] Here, the first terminal is equipped with a video capture device facing the user, such as a front-facing camera, which captures real-time user video.

[0132] Here, after an accessible call is established, the first terminal captures video facing the user.

[0133] It should be noted that, since hearing-impaired users cannot communicate verbally, in order to achieve the purpose of communication, during barrier-free calls, the hearing-impaired user, as the party in the call, expresses the content of the communication through body movements. The body movements of the hearing-impaired user are captured by the video capture device of the first terminal facing the user and generate real-time user video.

[0134] Understandably, in order to help hearing-impaired users confirm whether their communication is correct, the first terminal supports a local real-time preview function in the accessibility call mode. This allows hearing-impaired users to view their own body movement images in real time in the second area of ​​the second interface, and adjust their body language expression in a timely manner to ensure the accurate transmission of communication content.

[0135] In some embodiments, the method further includes: generating a first video stream based on a video facing the user; and sending the first video stream to a server.

[0136] It is understandable that the video captured by the first terminal facing the user is a video expressing the content of the user's communication. Therefore, after the first terminal captures the video facing the user, it encodes the video facing the user, and the first video stream generated after encoding is sent to the server as call data.

[0137] It is understandable that the first video stream contains the behavioral characteristics of the user on the first terminal side.

[0138] In some embodiments, since the user on the second terminal side is a hearing user who cannot recognize and understand the communication content expressed by the body movements of the user on the first terminal side, in the barrier-free call, after the server receives the first video stream, it decodes the first video stream and identifies the human behavior characteristics in the decoded first video stream, and converts the identified human behavior characteristics into voice content that the hearing user can understand, and then encodes the voice content into a first audio stream and sends it to the second terminal.

[0139] In some embodiments, the method further includes: displaying text information in a third area of ​​the second interface, the text information being used to characterize the communication content of the user of the second terminal.

[0140] Understandably, by displaying text information synchronously in the third area, hearing-impaired users can read the text subtitles while watching the sign language video in the first area, thus doubly confirming the communication intentions of the user on the other side and further improving the accuracy of information acquisition and communication efficiency.

[0141] In some embodiments, the method further includes: receiving text information sent by a server; and accordingly, displaying the text information in a third area of ​​the second interface, including: displaying the text information in the third area of ​​the second interface based on the text information.

[0142] Understandably, when the server converts the second audio stream into the second video stream in real time, it also generates corresponding text information based on the recognized speech content and sends the generated text information to the first terminal.

[0143] In one example, the first interface is as follows: Figure 5 As shown, after the user launches the accessibility calling application and completes user login, the first terminal's screen displays the first interface. The first interface displays icons for multiple terminals to be called, each consisting of a phone number icon and a call button icon. The user clicks the corresponding call button to indicate the terminal to be called, and the first terminal establishes an accessibility call with the corresponding terminal based on the user's click operation.

[0144] In one example, the second interface is as follows: Figure 6 As shown, the first area 601 is located in the middle of the second interface, the second area 602 is located in the upper right corner of the second interface, and the third area 603 is located at the bottom of the second interface. After the user clicks the call button, the first terminal and the second terminal establish an accessibility call. After the accessibility call is established, the screen display of the first terminal switches from the first interface to the second interface. Specifically, the first area 601 displays a sign language video of a virtual avatar, representing the communication content of the user on the other side; the second area 602 displays real-time body movement images of the user on the first terminal; and the third area 603 displays text subtitles expressing the same content as the sign language video of the virtual avatar.

[0145] In one application example of this application, such as Figure 7 As shown, the first terminal is specifically an accessible communication terminal, including a business module of an accessible calling application and a WebRTC module. The business module includes sub-modules such as login management, number management, dialing control, view management, and video rendering, which are used to start, execute, and terminate accessible calls. The WebRTC module includes sub-modules such as video capture, video encoding / decoding, dynamic bitrate adjustment, network assessment, and congestion control, which are used to establish a media channel and transmit video data between the first terminal and the server.

[0146] Specifically, the server is a cloud server. To implement the aforementioned barrier-free calling service function, the cloud server includes sub-modules such as business control, SIP gateway, WebRTC service, SIP service, AI gesture recognition, AI speech recognition, text-to-speech conversion, text-to-video conversion, video encoding / decoding, and audio encoding / decoding.

[0147] The call establishment method and communication method of this application embodiment will be described below in conjunction with the various functional sub-modules of the server described above.

[0148] In one application example of this application, an interaction method between a first terminal, a server, and a second terminal is provided, including a connection establishment phase and a communication transmission phase; such as Figure 8 As shown, the method includes: Step 801: The first terminal establishes a connection with the service control module of the server.

[0149] Here, after the first terminal starts the application, it connects to the server via the WebSocket protocol. The business control module manages the application login and acts as a signaling service.

[0150] Here, after logging into the first terminal application, the first terminal displays the first interface.

[0151] Here, WebSocket is a network protocol used to establish real-time bidirectional communication.

[0152] Step 802: The first terminal creates a PeerConnection object.

[0153] Here, before initiating a phone call, the first terminal, as the WebRTC initiator, first calls the Application Programming Interface (API) provided by WebRTC to create a PeerConnection object. The PeerConnection object is used to manage the media connection between the first terminal and the server, including session negotiation, network address traversal, and media stream transmission.

[0154] Step 803: The first terminal creates an Offer SDP based on the PeerConnection object.

[0155] Here, Offer SDP is the first request.

[0156] Step 804: The first terminal sends the Offer SDP to the business control module.

[0157] Here, the first terminal sends an Offer SDP based on the established WebSocket channel.

[0158] Step 805: The business control module forwards the Offer SDP to the WebRTC service module.

[0159] Here, the WebRTC service module is used to manage the establishment of media channels.

[0160] Step 806: The WebRTC service module creates the Answer SDP.

[0161] Here, the WebRTC service module fills the server's IP address into the Answer SDP as ICE Candidate information.

[0162] Step 807: The WebRTC service module sends the Answer SDP to the business control module.

[0163] Step 808: The service control module forwards the Answer SDP to the first terminal.

[0164] Here, the business control module sends Answer SDP based on the established WebSocket channel.

[0165] Step 809: The first terminal initiates a NAT Session Traversal Utilities for NAT Binding (STUN Binding) request to the WebRTC service module.

[0166] Here, the STUN protocol is a client-server protocol used to solve the NAT traversal problem. The STUNBinding request is used to obtain the public IP address and port after NAT translation. The server carries this public IP address in the response, enabling the requesting party to obtain its own public network mapping information, thereby providing an accessible network address for the subsequent establishment of media channels.

[0167] Step 810: The WebRTC service module responds to the STUN Binding request.

[0168] Here, the WebRTC service module sends a STUN Binding request response to the first terminal. The STUN Binding request response carries the public IP address of the first terminal.

[0169] Step 811: The first terminal establishes a media channel with the WebRTC service module.

[0170] Here, a media channel is established based on the public IP address of the first terminal and the IP address of the server, using the WebRTC protocol.

[0171] Step 812: The first terminal sends a call request to the service control module.

[0172] Here, the outgoing call request is used to request to establish an accessible call with the second terminal.

[0173] Here, the call-out request is the second request.

[0174] Step 813: The business control module forwards the outbound call request to the SIP service module.

[0175] Here, the SIP service module is used to manage carrier call setup.

[0176] Step 814: The SIP service module sends a SIP Invite signaling message to the SIP gateway module.

[0177] Here, the SIP gateway module is used to bridge the Internet and the operator's core network to achieve signaling interoperability.

[0178] Here, SIP Invite signaling is used to initiate a call, carrying media parameters (such as encoding format and network address) and requesting to establish a call with the second terminal.

[0179] Step 815: The SIP gateway module initiates a SIP Invite signaling to the second terminal.

[0180] Here, the SIP gateway module initiates SIP Invite signaling through the IMS network.

[0181] Step 816: The SIP gateway module sends a 100 Trying signal to the SIP service module.

[0182] Here, the 100 Trying signaling indicates that a request signaling is being processed.

[0183] Step 817: The second terminal sends a 100 Trying signaling message to the SIP gateway module.

[0184] Step 818: The second terminal starts ringing and sends a 180 Ringing signal to the SIP gateway module.

[0185] Here, the 180 Ringing signal indicates that the called end is ringing.

[0186] Step 819: The SIP gateway module forwards the 180 Ringing signaling to the SIP service module.

[0187] Step 820: The second terminal answers the call.

[0188] The call to be answered here is for accessible calls.

[0189] Step 821: The second terminal sends a 200 success (200 OK) signaling to the SIP gateway module.

[0190] Here, the 200 OK signaling is generated in response to the SIP Invite signaling, indicating that the request has been accepted.

[0191] Step 822: The SIP gateway module forwards the 200 OK signaling to the SIP service module.

[0192] In step 823, the SIP service module notifies the service control module of the listening event and sends an acknowledgment (ACK) signal to the SIP gateway module.

[0193] Step 824: The SIP gateway module forwards the ACK signaling to the second terminal.

[0194] Step 825: The SIP service module establishes a carrier call with the second terminal.

[0195] Step 826: The service control module sends a call answering notification to the first terminal.

[0196] Here, after the first terminal receives the call notification, it displays an interface indicating that the call is in progress.

[0197] Here, the interface representing the call is the second interface.

[0198] Here, the connection establishment phase includes steps 801 to 826 as described above.

[0199] Step 827: The first terminal acquires sign language video and performs video encoding to generate the first video stream.

[0200] Here, the first terminal uses a front-facing camera that is set to face the user to capture real-time user video, and obtains sign language video including sign language gesture information.

[0201] Here, the video encoding formats include, but are not limited to: H.264, VP8, and VP9. This example does not make any specific restrictions on these formats.

[0202] Step 828: The first terminal sends the first video stream to the WebRTC service module.

[0203] Here, the first terminal sends the first video stream based on the media channel established in step 811.

[0204] Step 829: The WebRTC service module decodes the first video stream.

[0205] Step 830: The WebRTC service module sends the decoded first video stream to the conversion module.

[0206] Here, the first video stream after decoding is a sign language video.

[0207] Here, the conversion module includes an AI gesture recognition module, an AI speech recognition module, a text-to-speech conversion module, and a text-to-video conversion module.

[0208] Step 831: The conversion module performs sign language recognition on the sign language video and generates text and voice information.

[0209] Here, the AI ​​gesture recognition module is configured with an AI behavior recognition model to perform sign language recognition on sign language videos and generate text information.

[0210] Here, the text-to-speech conversion module is used to convert text information into corresponding speech information.

[0211] Step 832: The conversion module sends the generated voice information to the SIP service module.

[0212] Step 833: The SIP service module performs audio encoding on the voice information to generate the first audio stream.

[0213] Step 834: The SIP service module sends the first audio stream to the second terminal.

[0214] Here, the SIP service module sends the first audio stream based on the carrier call established in step 825.

[0215] Step 835: The second terminal decodes the first audio stream.

[0216] Step 836: The second terminal plays the decoded audio data in real time.

[0217] Here, steps 827 to 836 are the communication transmission process from the first terminal to the second terminal.

[0218] Step 837: The second terminal collects voice information and performs audio encoding to generate a second audio stream.

[0219] Here, the second terminal collects the user's voice information based on the microphone.

[0220] Step 838: The second terminal sends the second audio stream to the SIP service module.

[0221] Here, the second terminal sends a second audio stream based on the operator call established in step 825.

[0222] Step 839: The SIP service module decodes the second audio stream.

[0223] Step 840: The SIP service module sends the decoded second audio stream to the conversion module.

[0224] Step 841: The conversion module recognizes the speech information and generates text information.

[0225] Here, the AI ​​speech recognition module is used to recognize speech information and generate text information.

[0226] Step 842: The conversion module generates a sign language video based on the text information.

[0227] Here, the text-to-video conversion module is configured with an AI behavioral feature generation model to generate corresponding sign language videos based on text information.

[0228] Here, the sign language video contains sign language gestures that the user can recognize.

[0229] Step 843: The conversion module sends text information to the business control module.

[0230] Step 844: The business control module sends text information to the first terminal.

[0231] Here, the business control module sends text messages based on the established WebSocket channel.

[0232] Step 845: The first terminal displays text information in real time.

[0233] Here, the first terminal displays text information in the third area of ​​the second interface.

[0234] Step 846: The conversion module sends the sign language video to the WebRTC service module.

[0235] Here, steps 846 and 843 are executed simultaneously.

[0236] Step 847: The WebRTC service module performs video encoding on the decoded sign language video to generate a second video stream.

[0237] Step 848: The WebRTC service module sends the second video stream to the first terminal.

[0238] Here, the WebRTC service module sends a second video stream based on the media channel established in step 811.

[0239] Step 849: The first terminal decodes the second video stream.

[0240] Step 850: The first terminal renders and displays the decoded second video stream in real time.

[0241] Here, the decoded second video stream is a sign language video.

[0242] Here, the first terminal displays the sign language video in the first area of ​​the second interface.

[0243] Here, steps 837 to 850 are the communication transmission process from the second terminal to the first terminal.

[0244] Here, the communication transmission stage includes steps 827 to 850 as described above.

[0245] In order to implement the method of the embodiments of this application, the embodiments of this application also provide a call establishment device for a server. The call establishment device corresponds to the call establishment method described above, and the steps in the embodiments of the call establishment method are also fully applicable to the embodiments of this device.

[0246] like Figure 9 As shown, the call establishment apparatus of this application embodiment includes one or more of the following modules: a first receiving module 901, a first identification module 902, a first sending module 903, a second receiving module 904, a second identification module 905, and a second sending module 906. Specifically, the first receiving module 901 receives a first video stream sent by a first terminal; the first identification module 902 identifies human behavior features in the first video stream and generates a first audio stream based on the identification result; the first sending module 903 sends the first audio stream to a second terminal based on a carrier call; the second receiving module 904 receives a second audio stream sent by the second terminal based on a carrier call; the second identification module 905 identifies the voice content in the second audio stream and generates a second video stream based on the identification result, the second video stream containing human behavior features used to characterize the voice content; and the second sending module 906 sends the second video stream to the first terminal.

[0247] In some embodiments, the call establishment apparatus includes a first receiving module 901, a first identification module 902, and a first sending module 903.

[0248] In some embodiments, the call establishment apparatus includes a second receiving module 904, a second identification module 905, and a second sending module 906.

[0249] In some embodiments, the call establishment device includes a first receiving module 901, a first identification module 902, a first sending module 903, a second receiving module 904, a second identification module 905, and a second sending module 906.

[0250] In some embodiments, the behavioral characteristics of a person include sign language gestures.

[0251] In some embodiments, the second recognition module 905 is further configured to generate corresponding text information based on the recognition result.

[0252] In some embodiments, the second sending module 906 is further configured to send text information to the first terminal.

[0253] In some embodiments, the first receiving module 901 is further configured to receive a first request sent by the first terminal; and establish a packet-switched network connection with the first terminal based on the first request; wherein the first video stream and / or the second video stream are transmitted based on the established packet-switched network connection.

[0254] In some embodiments, the first receiving module 901 is further configured to receive a second request sent by the first terminal, the second request carrying the identifier of the second terminal.

[0255] In some embodiments, the first sending module 903 is further configured to establish an operator call with the second terminal via a session initiation protocol based on a second request.

[0256] It should be noted that the call establishment device provided in the above embodiments is only illustrated by the division of the above-described program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the call establishment device and the call establishment method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0257] In order to implement the method of the embodiments of this application, the embodiments of this application also provide a communication device applied to a first terminal. The communication device corresponds to the above-described communication method, and the steps in the above-described communication method embodiments are also fully applicable to the embodiments of this device.

[0258] like Figure 10 As shown, the communication device in this embodiment includes a first display module 1001, a request module 1002, a second display module 1003, and a rendering module 1004. The first display module 1001 displays a first interface in an accessible call mode, the first interface including one or more icons representing terminals to be called; the request module 1002, in response to a user's selection of a target icon, requests the second terminal corresponding to the target icon to establish an accessible call; the second display module 1003 displays a second interface after the accessible call is established; the rendering module 1004 renders an image in a first area of ​​the second interface, the image displayed in the first area containing user-identifiable human behavioral characteristics, which represent the communication content of the user of the second terminal.

[0259] In some embodiments, the communication device further includes a data acquisition module 1005, which is used to acquire video directed toward the user.

[0260] In some embodiments, the rendering module 1004 is further configured to: render an image in a second area of ​​the second interface based on a video facing the user.

[0261] In some embodiments, the rendering module 1004 is further configured to: display text information in a third area of ​​the second interface, the text information being used to characterize the communication content of the user of the second terminal.

[0262] It should be noted that the communication device provided in the above embodiments is only illustrated by the division of the above program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the communication device and communication method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0263] Based on the hardware implementation of the above program modules, and in order to implement the call establishment method applied to the server in this application embodiment, this application embodiment also provides a server, such as... Figure 11 As shown, server 1100 includes at least one processor 1101, memory 1102, user interface 1103, and at least one network interface 1104. The various components in server 1100 are coupled together via bus system 1105. It can be understood that bus system 1105 is used to implement communication between these components. In addition to a data bus, bus system 1105 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 11 The general labeled all buses as Bus System 1105.

[0264] The user interface 1103 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0265] The memory 1102 in this embodiment is used to store various types of data to support the operation of the server 1100. Examples of such data include any computer program used to operate on the server 1100.

[0266] The call establishment method disclosed in this application can be applied to or implemented by the processor 1101. The processor 1101 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the call establishment method can be completed by integrated logic circuits in the hardware of the processor 1101 or by instructions in software form. The processor 1101 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1101 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in a storage medium, which is located in the memory 1102. The processor 1101 reads information from the memory 1102 and, in conjunction with its hardware, completes the steps of the call establishment method provided in the embodiments of this application.

[0267] In an exemplary embodiment, server 1100 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned call establishment method.

[0268] It is understood that memory 1102 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), EEPROM, ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM). The memory 1102 described in this application embodiment is intended to include, but is not limited to, these and any other suitable types of memory.

[0269] Based on the hardware implementation of the above program modules, and in order to implement the communication method applied to the first terminal in this application embodiment, this application embodiment also provides a first terminal, such as... Figure 12 As shown, the first terminal 1200 includes at least one processor 1201, a memory 1202, a user interface 1203, and at least one network interface 1204. The various components in the first terminal 1200 are coupled together via a bus system 1205. It can be understood that the bus system 1205 is used to implement communication between these components. In addition to a data bus, the bus system 1205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 12 The general labeled all buses as Bus System 1205.

[0270] The user interface 1203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0271] The memory 1202 in this embodiment is used to store various types of data to support the operation of the first terminal 1200. Examples of such data include any computer program used to operate on the first terminal 1200.

[0272] The communication method disclosed in this application embodiment can be applied to or implemented by the processor 1201. The processor 1201 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the communication method can be completed by the integrated logic circuit of the hardware in the processor 1201 or by instructions in the form of software. The processor 1201 mentioned above may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1201 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 1202. The processor 1201 reads the information in the memory 1202 and, in conjunction with its hardware, completes the steps of the communication method provided in the embodiments of this application.

[0273] In an exemplary embodiment, the first terminal 1200 may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned communication method.

[0274] It is understood that memory 1202 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be ROM, PROM, EPROM, EEPROM, FRAM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM; magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be RAM, which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as SRAM, SSRAM, DRAM, SDRAM, DDRSDRAM, ESDRAM, SLDRAM, and DRRAM. The memory 1202 described in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0275] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 1102 including a computer program, which can be executed by the processor 1101 of the server 1100 to complete the steps described in the call establishment method of this application embodiment; or, a memory 1202 including a computer program, which can be executed by the processor 1201 of the first terminal 1200 to complete the steps described in the communication method of this application embodiment. The computer-readable storage medium can be a ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.

[0276] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by the processor 1101 of the server 1100 to complete the steps described in the method of this application embodiment, or executed by the processor 1201 of the first terminal 1200 to complete the steps described in the method of this application embodiment.

[0277] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0278] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0279] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for establishing a call, characterized in that, Applied to a server, the method includes: Receive the first video stream sent by the first terminal; Identify the behavioral characteristics of people in the first video stream, and generate a first audio stream based on the identification results; Based on the carrier call, the first audio stream is sent to the second terminal; And / or, Based on the operator's call, receive the second audio stream sent by the second terminal; The speech content in the second audio stream is identified, and a second video stream is generated based on the identification result. The second video stream contains human behavioral features used to characterize the speech content. Send the second video stream to the first terminal.

2. The method according to claim 1, characterized in that, The behavioral characteristics of the person mentioned include sign language gestures.

3. The method according to claim 1, characterized in that, After recognizing the speech content in the second audio stream, the method further includes: Generate corresponding text information based on the recognition results; The text information is sent to the first terminal.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Receive the first request sent by the first terminal; Based on the first request, establish a packet-switched network connection with the first terminal; The first video stream and / or the second video stream are transmitted based on an established packet-switched network connection.

5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Receive a second request sent by the first terminal, wherein the second request carries the identifier of the second terminal; Based on the second request, a call with the operator is established with the second terminal via a session initiation protocol.

6. A communication method, characterized in that, Applied to a first terminal, the method includes: In the accessibility call mode, a first interface is displayed, which includes one or more icons representing the terminal to be called; In response to the user's selection of a target icon, a request is made to establish an accessibility call with the second terminal corresponding to the target icon; After the barrier-free call is established, the second interface is displayed; An image is rendered in a first area of ​​the second interface. The image displayed in the first area contains user-identifiable human behavioral characteristics, which characterize the communication content of the user of the second terminal.

7. The method according to claim 6, characterized in that, The method further includes: Capture video facing the user; Based on the video facing the user, an image is rendered in the second area of ​​the second interface.

8. The method according to claim 6, characterized in that, The method further includes: Text information is displayed in the third area of ​​the second interface, and the text information is used to represent the communication content of the user of the second terminal.

9. A server, characterized in that, The server includes: A processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method according to any one of claims 1 to 5.

10. A first terminal, characterized in that, The first terminal includes: A processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method according to any one of claims 6 to 8.