Call method and device, network equipment, electronic equipment and storage medium
By utilizing information input on the dial pad during a call, combined with large-scale model recognition and conversion processing, the inconvenience of voice communication and the problem of face exposure in video calls are solved, enabling more convenient information transmission and display, and improving the convenience of calls.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE INTERNET CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-19
AI Technical Summary
During the call, inconvenient voice communication, inability to send text messages, the need to show one's face for video calls, and inaccurate gesture communication all contribute to poor call convenience.
By inputting information on the dial pad, a large model is used for recognition and conversion processing to obtain a set of text information, which is then converted into media information and sent to the other end device for display, supporting display in the form of voice, video, or images.
It improves the convenience of calls, reduces the inconvenience of voice communication, the inability to send text messages, and the problem of face exposure in video calls, and enables more complex information exchange and real-time rendering.
Smart Images

Figure CN122069320A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a method of making a call, an apparatus, a network device, an electronic device, and a storage medium. Background Technology
[0002] With the development of science and technology, the convenience of communication has become a major focus. During calls, if voice communication is inconvenient, text messages or video calls can be used. However, text messaging suffers from delays and the inability to send messages to the other end. Video calls require gestures, and the inability to understand gestures reduces the convenience of the call. Summary of the Invention
[0003] This disclosure provides a call method, apparatus, network device, electronic device, and storage medium that can improve the convenience of calls. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a call method is provided, comprising: Receive first input information sent by a first electronic device that is in the process of a call, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; The first input information is identified and converted to obtain a set of text information corresponding to the first input information. Send the set of text information to the first electronic device; According to the selection instruction sent by the first electronic device for the set of text information, the target text information is obtained; The target text information is converted to obtain the media information corresponding to the target text information; The media information is sent to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the first electronic device and the second electronic device are devices for conducting a call.
[0004] According to some embodiments, the method further includes: When the first electronic device initiates a call, or before the first electronic device initiates a call, the system negotiates with the first electronic device to obtain the display method corresponding to the text information set, wherein the display method is used to indicate the video format for displaying the text information set.
[0005] According to some embodiments, obtaining the set of text information corresponding to the first input information includes: The first input information is used to draw an image on the canvas, and the image is decoded to obtain the first video data in the video format; During the process of acquiring the first video data, if no second input information is received from the first electronic device, the first real-time transmission protocol packet corresponding to the first video data is acquired and sent to the first electronic device.
[0006] According to some embodiments, the method further includes: During the process of acquiring the first video data, upon receiving the second input information sent by the first electronic device, the first video data is updated to acquire the second video data; Obtain the second real-time transmission protocol packet corresponding to the second video data, and send the second real-time transmission protocol packet to the first electronic device.
[0007] According to some embodiments, the media information includes voice information, and sending the media information to the second electronic device includes: If the format of the voice information does not conform to the encoding negotiated by the communication signaling, the voice information is decoded into voice information in Pulse Code Modulation (PCM) format. The PCM format speech information is re-encoded into a negotiated encoding format speech information; The negotiated encoded voice information is sent to the second electronic device.
[0008] According to a second aspect of the present disclosure, a call method is provided, applied to a first electronic device, comprising: While the first electronic device is in the middle of a call, the first input information is received on the dial pad of the first electronic device; Send the first input information to the network device, wherein the first input information is used to instruct the network to use a large model to recognize and convert the first input information to obtain the set of text information corresponding to the first input information; Receive and display a set of text information sent by a network device, and receive a selection instruction corresponding to the set of text information; The selection instruction is sent to the network device, wherein the selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the second electronic device is the peer device in the call process.
[0009] According to some embodiments, the method further includes: If the first input information is not received on the dial pad of the first electronic device, preset video information is displayed on the screen of the first electronic device.
[0010] According to some embodiments, the method further includes: Obtain the first time point of the first button on the dial pad, wherein the first button is the start button on the dial pad; Obtain the second time point of the second button on the dial pad, wherein the second button is the end button on the dial pad; Based on the first time point and the second time point, obtain the instruction type corresponding to the input instruction; Based on the instruction type, obtain the input instruction corresponding to the dial pad.
[0011] According to some embodiments, obtaining the instruction type corresponding to the input instruction based on the first time point and the second time point includes: Based on the first time point and the second time point, obtain the duration between the first time point and the second time point; If the duration is less than the first duration threshold, the instruction type corresponding to the input instruction is obtained as the input instruction type; If the duration is greater than the second duration threshold, the instruction type corresponding to the input instruction is determined to be a control instruction type, wherein the second duration threshold is greater than or equal to the first duration threshold.
[0012] According to some embodiments, the set of text information sent by the network device includes: When displaying the set of text information and receiving a trigger command for a preset key on the dial pad, the set of text information is displayed using a display method corresponding to the preset key.
[0013] According to a third aspect of the present disclosure, a call method is provided, applied to a second electronic device, comprising: The network device receives media information sent by a network device, wherein the media information is media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set, and the target text information is converted and processed. Display the aforementioned media information.
[0014] According to a fourth aspect of the present disclosure, a communication device is provided, comprising: The first information receiving unit is configured to receive first input information sent by a first electronic device during a call, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; The set acquisition unit is used to perform recognition and conversion processing on the first input information to acquire the set of text information corresponding to the first input information; A collection sending unit is used to send the collection of text information to the first electronic device; The information acquisition unit is used to acquire target text information according to the selection instruction sent by the first electronic device for the text information set; The information acquisition unit is also used to convert the target text information to obtain the media information corresponding to the target text information; An information sending unit is used to send the media information to a second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the first electronic device and the second electronic device are devices for conducting a call.
[0015] According to a fifth aspect of the present disclosure, a communication device is provided, comprising: The second information receiving unit is used to receive first input information on the dial pad of the first electronic device when the first electronic device is in a call. A collection receiving unit is used to receive and display a collection of text information sent by a network device, and to receive a selection instruction corresponding to the collection of text information, wherein the collection of text information is a collection of information obtained by the network device through recognition processing and conversion processing based on the first input information sent by the first electronic device; The instruction sending unit is used to send the selection instruction to the network device, wherein the selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the second electronic device is the peer device in the call process.
[0016] According to a sixth aspect of the present disclosure, a communication device is provided, comprising: The third information receiving unit is used to receive media information sent by the network device, wherein the media information is media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set, and the target text information is converted and processed. An information display unit is used to display the media information.
[0017] According to a seventh aspect of the present disclosure, an electronic device is provided, comprising: A processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the call method described in any one of the preceding aspects.
[0018] According to an eighth aspect of the present disclosure, a storage medium is provided that, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform any of the preceding aspects of the call method.
[0019] According to a ninth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described in any one of the preceding aspects.
[0020] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In some or related embodiments, the process involves receiving first input information sent by a first electronic device during a call, wherein the first input information is received by the first electronic device on its dial pad; performing recognition and conversion processing on the first input information to obtain a set of text information corresponding to the first input information; sending the set of text information to the first electronic device; obtaining target text information according to a selection instruction sent by the first electronic device for the set of text information; performing conversion processing on the target text information to obtain media information corresponding to the target text information; and sending the media information to a second electronic device, wherein the media information is used to instruct the second electronic device to display the media information. The first electronic device and the second electronic device are devices conducting the call. Therefore, information can be input via the dial pad on an electronic device during a call, providing a call method that reduces the inconvenience of voice communication, reduces the situation where SMS messages cannot be sent to the other end, reduces the need for face recognition and inaccurate gestures in video calls, reduces the situation where the other end cannot understand the content to be communicated, and reduces the situation where the preset voice method does not match the current content to be expressed. This allows for more complex calls, real-time rendering of input information, conversion of text into media information, and transmission to the other end, thus improving the convenience of the call.
[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0023] Figure 1 This is a flowchart of the first call method provided in the embodiments of this disclosure; Figure 2 A schematic diagram illustrating an example of a call method according to an embodiment of this disclosure is shown; Figure 3A This is an example schematic diagram of the first type of electronic device display interface provided in the embodiments of this disclosure; Figure 3B This is an example schematic diagram of a second type of electronic device display interface provided in this embodiment of the present disclosure; Figure 4 This is a flowchart illustrating a second call method according to an embodiment of the present disclosure; Figure 5 This is a flowchart illustrating a third type of call method according to an embodiment of the present disclosure; Figure 6 This is a block diagram illustrating a first type of communication device according to an exemplary embodiment; Figure 7 This is a block diagram illustrating a second type of communication device according to an exemplary embodiment; Figure 8 This is a block diagram illustrating a third type of communication device according to an exemplary embodiment; Figure 9 This is an example schematic diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0025] This disclosure provides embodiments of a communication method, apparatus, electronic device, and storage medium. In some embodiments, the terms "communication method" and "information processing method," "communication method," etc., can be used interchangeably; the terms "communication apparatus" and "information processing apparatus," "communication apparatus," etc., can be used interchangeably; and the terms "information processing system," "communication system," etc., can be used interchangeably.
[0026] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0027] In all embodiments disclosed herein, unless otherwise specified or logically conflicting, the terminology and / or descriptions are consistent across embodiments and can be referenced interchangeably. Technical features from different embodiments can be combined to form new embodiments based on their inherent logical relationships. The terminology used in the embodiments of this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of this disclosure.
[0028] In this disclosure, unless otherwise stated, elements expressed in the singular, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular or a plural expression. In this disclosure, "multiple" refers to two or more. In some embodiments, terms such as "at least one of," "one or more," "a plurality of," and "multiple" can be used interchangeably.
[0029] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0030] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0031] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0032] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0033] Figure 1 This is a flowchart of the first call method provided in the embodiments of this disclosure, such as... Figure 1 As shown, it includes the following steps: In step S11, the first input information sent by the first electronic device in the middle of a call is received, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; According to some embodiments, the implementing entity of this disclosure may be, for example, a network device. This network device does not specifically refer to a particular fixed network device. For example, when the device identifier corresponding to the network device changes, the network device may also change accordingly. For example, when the device structure of the network device changes, the network device may also change accordingly.
[0034] In some embodiments, the call method can be used in call scenarios where it is inconvenient to speak, such as in daily life, where there are situations where people with disabilities cannot speak normally when answering the phone due to throat diseases, or when some people with speech impairments cannot communicate normally by voice.
[0035] According to some embodiments, a call process can be, for example, the process of a calling number initiating a call and the call starting after the call is successful. This call process is not specifically defined by a fixed procedure. For example, the call process can change accordingly when the parties involved in the call change. Similarly, the call process can change accordingly when the time point corresponding to the call changes.
[0036] In some embodiments, the process may be a call between a first electronic device and a second electronic device. The first electronic device can be any one of the electronic devices involved in the call. This first electronic device does not specifically refer to a fixed device. For example, if the device identifier of the first electronic device changes, the first electronic device may also change accordingly. For example, if the device structure of the first electronic device changes, the first electronic device may also change accordingly. The electronic device may also be referred to as a terminal, terminal device, mobile terminal, etc. The first electronic device may be a service subscription terminal, that is, a terminal that has subscribed to text input and text conversion functions via a dial pad.
[0037] According to some embodiments, the first input information may be, for example, input information received by the first terminal device on the dial pad of the first terminal device during a call. This first input information does not specifically refer to any fixed information. For example, when the numerical information corresponding to the first input information changes, the first input information may also change accordingly. For example, when the input time of the first input information changes, the first input information may also change accordingly.
[0038] In some embodiments, the dial pad may be, for example, a dial pad used by the first terminal device for dialing.
[0039] In some embodiments, a first input message sent by a first electronic device during a call can be received, wherein the first input message is the first input message received by the first electronic device on the dial pad of the first electronic device.
[0040] According to some embodiments, the method further includes: When the first electronic device initiates a call, or before the first electronic device initiates a call, the system negotiates with the first electronic device to obtain the display method corresponding to the text information set, wherein the display method is used to indicate the video format for displaying the text information set.
[0041] In some embodiments, the display mode can be used to indicate how a set of text information is displayed. For example, when the display mode is a video display mode, the set of text information can be displayed as video information. The display mode does not specifically refer to any particular fixed information. For example, the display mode can also change when the negotiation method or negotiation time changes.
[0042] There is no limitation on the timing of determining the display method. For example, it can be determined during the current call or in advance.
[0043] According to some embodiments, Figure 2 A schematic diagram illustrating an example of a call method according to an embodiment of this disclosure is shown. For example... Figure 2 As shown, the call system may include a terminal module (electronic device), a signaling module (AS), a core media service module, a unified media processing module, a media rendering module, and an input method large model module. Among them: Terminal module: The mobile terminal through which users interact directly. The terminal's dial interface has a dial pad, and this solution operates on the dial pad.
[0044] Signaling Module (Application Server, AS): Performs media negotiation on incoming signaling, including audio, video, and digit collection, and controls the media of dual-number calls to the corresponding media module.
[0045] Core Media Server (MEDIA): Responsible for media stream forwarding, packaging media data into Real-time Transport Protocol (RTP) data, and playing audio to mobile terminals, etc.
[0046] Unified Media Function (UMF): Responsible for text-to-speech (TTS) conversion, implementing speech encoding and decoding, and acting as an intermediary service connecting various media layers.
[0047] Video rendering service: It can render text onto a canvas in real time, generate images, and encode images into videos in real time.
[0048] The input method model displays a list of characters based on the nine-key input method, essentially a cloud-based input method.
[0049] In some embodiments, the display method may be determined after the terminal initiates a call, i.e., the display method of the information is determined after the terminal initiates a call. Specifically, the signaling module may start the Session Description Protocol (SDP) negotiation and enable video function and Dual-Tone Multi-Frequency (DTMF) number collection function for the service subscriber.
[0050] The subscription terminal enables video functionality. During SDP negotiation, voice calls are prioritized for AMR encoding, while video calls are negotiated for H.264 format. On terminals with the subscription function, both voice and video ports are negotiated; on terminals without the subscription function, only audio functionality is enabled.
[0051] In signaling negotiation, AMR encoding is prioritized for audio, and the SDP carries a telephone event to initiate 2833 digit collection. The Real-time Transport Protocol (RTP) can also include 2833 digit collection events. In the core media service, the 2833 digit collection length is set to be unlimited, and a timeout is also set, allowing continuous digit collection. If G711 encoding is used, the G algorithm (Goertzel) can also be used to obtain the digit collection event; this disclosure does not limit this approach.
[0052] When negotiating media videos, high-definition (720P) standards are prioritized. If negotiation fails, the video is downgraded to standard definition (480P).
[0053] According to some embodiments, on the dial pad interface of the first electronic device, typing can be performed using the T9 input method, and the core media service can be reached by detecting the events of the pressed numbers and DTMF events; wherein, RFC2833 specifies the standard for transmitting DTMF numbers and other telephone tones and signals. The payload format of RFC2833 can be, for example, as shown in Table 1.
[0054] Table 1
[0055] Specifically, for the first electronic device, a 2833 sequence packet can be, for example, from pressing a key to releasing it. Multiple sequence packets may occur during this period, with the RTP sequence packet timestamps from the start to the end being consistent. The duration increases, allowing the system to determine whether it's a long press or a short press. A short press could be, for example, an input event in an input method, while a long press could be, for example, a control event. Specifically, while continuously pressing keys to input text, a long press can be used to select a sentence, or *# can be used to turn pages.
[0056] In step S12, the first input information is identified and converted to obtain the set of text information corresponding to the first input information; According to some embodiments, a large model can be used to recognize and transform the first input information to obtain a set of text information corresponding to the first input information. The large model can also be called an input method large model, and it does not specifically refer to a fixed model. For example, when the model type corresponding to the large model changes, the large model can also change accordingly. For example, when the parameters in the large model change, the large model can also change accordingly.
[0057] In some embodiments, the text information set may be, for example, a collection of at least one text information obtained by processing the first input information using a large model. This text information set does not specifically refer to a fixed set. For example, the text information set may change accordingly when the number of text information items included in it changes. Similarly, the text information set may change accordingly when a particular text information item included in it changes.
[0058] According to some embodiments, the first input information can be identified and converted to obtain a set of text information corresponding to the first input information.
[0059] In step S13, a set of text messages is sent to the first electronic device; According to some embodiments, when a set of text information is obtained, the set of text information can be sent to a first electronic device. When the first electronic device receives the set of text information, it can display the set of text information.
[0060] In step S14, the target text information is obtained according to the selection instruction sent by the first electronic device for the text information set; According to some embodiments, the selection instruction may be, for example, a selection instruction input to a set of text information received by a first electronic device. This selection instruction includes, but is not limited to, click selection instructions, voice selection instructions, etc. The selection instruction is not specifically a fixed instruction. For example, when the receiving time point corresponding to the selection instruction changes, the selection instruction may also change accordingly.
[0061] In some embodiments, the target text information can be, for example, a certain piece of information selected from a set of text information. The target text information can be, for example, the information corresponding to a selection instruction. The name of the target text information is not limited. For example, the target text information can also be referred to as the text information to be converted, the selected text information, the first text information, etc. The target text information does not specifically refer to a certain fixed information. For example, when the selection instruction changes, the target text information can also change accordingly. For example, when the set of text information changes, the target text information can also change accordingly.
[0062] According to some embodiments, the target text information can be obtained based on a selection instruction sent by a first electronic device for a set of text information.
[0063] According to some embodiments, as Figure 3A shown, when "94664" is entered on the dial pad based on T9, that is, the first input information is "94664", and the corresponding pinyin is "zhong", and there are four Chinese character choices, namely 中 (zhong), 钟 (zhong), 重 (zhong), and 种 (zhong), that is, the set of text information can include 中 (zhong), 钟 (zhong), 重 (zhong), and 种 (zhong). Continuing to enter "94" on the dial pad of the first electronic device, that is, the second input information is "94", and the obtained pinyin becomes "zhongyi", and there are multiple groups of corresponding Chinese words in the updated set of text information, for example, it can be as Figure 3B shown. Among them, if the user of the first electronic device needs to select "中移" (China Mobile), then long press "1", and the first electronic device can receive the selection instruction and send it to the network device, and the network device can determine in the core media service that the word to be selected is the 1st "中移" (China Mobile).
[0064] Specifically, as Figure 2 shown, the core media service transmits all DTMF call receiving events to the "Unified Media Service", and the Unified Media Service records the call receiving events and transmits the call receiving events to the input method large model each time. Every time the large model receives a number, it returns a list of phrases to the Unified Media Service based on the nine - grid input method of Text on 9 keys (T9); cleaning events, etc. are defined between the Unified Media Service and the input method large model. When the font selection is completed, the cache is cleared, and thus the conversion from numbers to text is completed.
[0065] In step S15, perform conversion processing on the target text information to obtain the media information corresponding to the target text information; According to some embodiments, the conversion process for the target text information may include format conversion of the target text information, such as converting text information into voice information or video information. The conversion process in this disclosure is not specifically defined by a single fixed process. For example, the conversion process may change accordingly when the target text information changes. Similarly, the conversion process may change accordingly when the way media information is displayed changes.
[0066] According to some embodiments, the name of the media information is not limited. The media information may also be referred to as multimedia information. This multimedia information includes, but is not limited to, image information, video information, and audio information.
[0067] In some embodiments, the target text information can be converted to obtain the media information corresponding to the target text information.
[0068] According to some embodiments, sending a collection of text messages to a first electronic device includes: The text information is collected on the canvas to draw an image, and the image is decoded to obtain the first video data in video format; During the acquisition of the first video data, if no second input information is received from the first electronic device, the first real-time transmission protocol packet corresponding to the first video data is acquired and sent to the first electronic device. Therefore, the text information set can be rendered, and the corresponding video data can be acquired, thus improving the accuracy of video data acquisition.
[0069] According to some embodiments, the first video data may be, for example, video data determined based on the first input information. This first video data does not specifically refer to any fixed data. For example, when the target text information changes, the target text information may also change accordingly. For example, when the method of generating the first video data changes, the target text information may also change accordingly.
[0070] In some embodiments, the second input information may be information entered after the first input information. The number of second input information entries may be multiple or a single entry. The second input information in this disclosure does not specifically refer to any fixed information. For example, if the input time of the second input information changes, the second input information may also change accordingly. For example, if the content of the second input information changes, the second input information may also change accordingly.
[0071] In some embodiments, the first real-time transport protocol packet may be, for example, a carrier for transmitting the first video data. This first real-time transport protocol packet does not specifically refer to a particular fixed real-time transport protocol packet. For example, when the method of generating the real-time transport protocol packet changes, the first real-time transport protocol packet may also change accordingly. For example, when the first video data changes, the first real-time transport protocol packet may also change accordingly.
[0072] According to some embodiments, the method further includes: During the process of acquiring the first video data, upon receiving the second input information sent by the first electronic device, the first video data is updated and the second video data is acquired. Obtain the second real-time transmission protocol packet corresponding to the second video data, and send the second real-time transmission protocol packet to the first electronic device.
[0073] According to some embodiments, the method further includes: If the format of the media information does not conform to the encoding negotiated by the communication signaling, the media information will be decoded into PCM format media information; Re-encode PCM format media information into a negotiated encoding format.
[0074] According to some embodiments, when the Unified Media Service receives the WordList, it can transmit the WordList to the Media Rendering Service. The Media Rendering Service can draw the WordList as an image on a canvas, decode the image to obtain Luminance-Bandwidth-Chrominance (YUV) data, and then encode the YUV data into H.264 data in real time to obtain the first video data. During this process, if there is no new WordList, the old image is rendered. When the canvas completes the drawing of a new image, the new image is rendered. Video rendering is continuous, and the rendering is based on the YUV data obtained from the buffer image decoding, which is encoded at intervals. Video rendering can be based on the H.264 video ringback tone standard, for example, and continues to render according to the negotiated display method, which can include 720P, 480P, ultra-high definition 1080P, etc. The rendering frame rate can be, for example, 30 frames per second. The rendered video is transmitted to the Unified Media Service through a real-time interface, and then the Unified Media Service transmits it back to the core media service.
[0075] According to some embodiments, the rendered video (H264 format) arrives at the core media service, where a single frame of H264 is packaged into multiple RTP packets with the same time sequence, and then sent to the service subscription terminal. During the packaging process, the H264 data is transmitted to the core media service frame by frame, and the video information is contained in the Picture Parameter Set (PPS) and Sequence Parameter Set (SPS). H264 is split into multiple RTP packets based on a single frame of data, and the maximum transmission unit (MTU) of each RTP is limited to no more than 1420.
[0076] When the data packet arrives at the terminal (the first electronic device), a video is displayed on the dialer interface. The video shows the input keys, followed by a list of phrases. A long press allows selection of the chosen text. This sends a selection command to the network device, which then determines the target text. The selected text is determined within the Unified Media Service.
[0077] In step S16, media information is sent to the second electronic device, wherein the media information is used to instruct the second electronic device to display media information, and the first electronic device and the second electronic device are devices for conducting a call.
[0078] According to some embodiments, the second electronic device may be an electronic device that communicates with the first electronic device. The second electronic device does not specifically refer to a particular fixed device. For example, if the device identifier of the second electronic device changes, the second electronic device may also change accordingly.
[0079] According to some embodiments, wherein the media information includes voice information, sending the media information to the second electronic device includes: If the format of the voice information does not conform to the encoding negotiated by the communication signaling, the voice information will be decoded into PCM format voice information. The PCM format speech information is re-encoded into a negotiated encoding format speech information; Send voice information in the negotiated encoding format to the second electronic device.
[0080] In some embodiments, such as Figure 2As shown, the selected text is converted into voice data via the TTS interface. Since the current voice data does not conform to the encoding negotiated in the communication signaling agreement, it needs to be decoded into PCM first, and then re-encoded into the currently negotiated encoding format. The core media notifies the unified media service of the current audio encoding format. The encoded and decoded voice can be transmitted to the core media service, which will package it into RTP packets and play them on unsubscribed service terminals so that they can hear the content of the text selected by typing in the "subscribed service terminal". During playback, RTP packets sent from the "subscribed service terminal" will not be forwarded to the "unsubscribed service terminal," forcing the use of the text-to-speech (TTS)-derived voice for playback.
[0081] In some or related embodiments, the process involves receiving first input information sent by a first electronic device during a call, wherein the first input information is received by the first electronic device on its dial pad; performing recognition and conversion processing on the first input information to obtain a set of text information corresponding to the first input information; sending the set of text information to the first electronic device; obtaining target text information according to a selection instruction sent by the first electronic device for the set of text information; performing conversion processing on the target text information to obtain media information corresponding to the target text information; and sending the media information to a second electronic device, wherein the media information is used to instruct the second electronic device to display media information. The first and second electronic devices are devices used for the call. Therefore, information can be input via the dial pad on an electronic device during a call, providing a call method that reduces the inconvenience of voice communication, reduces the situation where SMS messages cannot be sent to the other end, reduces the need for face recognition and inaccurate gestures in video calls, reduces the situation where the other end cannot understand the content to be communicated, and reduces the situation where the preset voice method does not match the current content to be expressed. This allows for more complex calls, real-time rendering of input information, conversion of text into media information, and transmission to the other end, thus improving the convenience of the call. Alternatively, you can use the DTMF function of the 2833 to collect numbers, combine it with the nine-grid input method to achieve text input, then use the video function to demonstrate the text input effect and select text, and finally convert the text to speech via TTS and force the audio to be played on the other end to achieve voice communication.
[0082] According to some embodiments, a call method is also provided, applied to a first electronic device, comprising: While the first electronic device is in the middle of a call, the first input information is received on the dial pad of the first electronic device; Receive and display a set of text information sent by a network device, and receive a selection instruction corresponding to the set of text information. The set of text information is a set of information obtained by the network device through recognition, processing and conversion of the first input information sent by the first electronic device. A selection instruction is sent to the network device, wherein the selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, convert the target text information, obtain the media information corresponding to the target text information, and send the media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the second electronic device is the peer device in the call process.
[0083] Figure 4 This is a flowchart of the second call method provided in the embodiments of this disclosure, such as... Figure 4 As shown, it includes the following steps: In step S21, while the first electronic device is in a call, the first input information is received on the dial pad of the first electronic device; The relevant processes can be as described above, and will not be repeated here.
[0084] According to some embodiments, the first electronic device may be, for example, a service subscription terminal. The second electronic device may be, for example, a non-subscribed service terminal.
[0085] According to some embodiments, the method further includes: If no first input information is received on the dial pad of the first electronic device, preset video information is displayed on the screen of the first electronic device.
[0086] According to some embodiments, such as Figure 2 As shown, the method may further include: the core media service in the network device can act as a proxy connecting the calling terminal and the called terminal, and can transparently transmit the calling voice RTP to the called terminal and the called RTP to the calling terminal, thus enabling voice communication between the two parties. Specifically, for terminals subscribed to the service, video RTP can be enabled, and the media rendering service transmits the media stream to the core media service. The core media service loops and plays the default video data on the terminals subscribed to the service. At this time, the voice RTP starts listening for DTMF events. If a button is pressed on the dial pad by the terminal subscribed to the service, the corresponding event is transmitted to the core media service.
[0087] According to some embodiments, the method further includes: Get the first time point of the first key on the dial pad, where the first key is the start key on the dial pad; Get the second time point of the second button on the dial pad, where the second button is the end button on the dial pad; Based on the first and second time points, obtain the instruction type corresponding to the input instruction; Based on the instruction type, retrieve the input instruction corresponding to the dial pad.
[0088] According to some embodiments, the instruction type corresponding to the input instruction is obtained based on a first time point and a second time point, including: Based on the first and second time points, obtain the duration between the first and second time points; If the duration is less than the first duration threshold, the instruction type corresponding to the input instruction is obtained as the input instruction type; If the duration is greater than the second duration threshold, the instruction type corresponding to the input instruction is obtained as a control instruction type, wherein the second duration threshold is greater than or equal to the first duration threshold.
[0089] In step S22, the first input information is sent to the network device, wherein the first input information is used to instruct the network to use a large model to recognize and convert the first input information to obtain the set of text information corresponding to the first input information; The relevant processes can be as described above, and will not be repeated here.
[0090] In step S23, the set of text information sent by the network device is received and displayed, and the selection instruction corresponding to the set of text information is received; The relevant processes can be as described above, and will not be repeated here.
[0091] According to some embodiments, a set of text messages sent by a network device is shown, including: When displaying a set of text information and receiving a trigger command for a preset key on the dial pad, the set of text information is displayed using a display method corresponding to the preset key.
[0092] According to some embodiments, for example, when the first input information corresponds to multiple word list options, the corresponding number button can be long-pressed to select the corresponding item. If page turning is required, the "*" button is the previous page, and pressing "#" is the next page.
[0093] In step S24, a selection instruction is sent to the network device. The selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device. The media information is used to instruct the second electronic device to display the media information. The second electronic device is the peer device in the call process.
[0094] The relevant processes can be as described above, and will not be repeated here.
[0095] Figure 5 This is a flowchart of the third call method provided in the embodiments of this disclosure, such as... Figure 5 As shown, it includes the following steps: In step S31, media information sent by the network device is received. The media information is the media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set and the conversion processing of the target text information. The relevant processes can be as described above, and will not be repeated here.
[0096] In step S32, media information is displayed.
[0097] The relevant processes can be as described above, and will not be repeated here.
[0098] The media information can be audio information.
[0099] Figure 6 This is a schematic diagram of the first type of communication device provided in the embodiments of this disclosure, such as... Figure 6 As shown, it includes: The first information receiving unit 601 is used to receive first input information sent by the first electronic device during a call, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; The set acquisition unit 602 is used to perform recognition and conversion processing on the first input information to acquire the set of text information corresponding to the first input information; The collection sending unit 603 is used to send a collection of text information to the first electronic device; The information acquisition unit 604 is used to acquire target text information according to the selection instruction sent by the first electronic device for the text information set; The information acquisition unit 604 is also used to convert and process the target text information to acquire the media information corresponding to the target text information; The information sending unit 605 is used to send media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the first electronic device and the second electronic device are devices for conducting a call.
[0100] According to some embodiments, the first information receiving unit 601 is further configured to: When the first electronic device initiates a call, or before the first electronic device initiates a call, the system negotiates with the first electronic device to obtain the display method corresponding to the text information set, wherein the display method is used to indicate the video format for displaying the text information set.
[0101] According to some embodiments, when the information acquisition unit 604 sends a set of text information to the first electronic device, it is specifically used for: The text information is collected on the canvas to draw an image, and the image is decoded to obtain the first video data in video format; During the acquisition of the first video data, if the second input information sent by the first electronic device is not received, the first real-time transmission protocol packet corresponding to the first video data is acquired and sent to the first electronic device.
[0102] According to some embodiments, the information acquisition unit 604 is further configured to: During the process of acquiring the first video data, upon receiving the second input information sent by the first electronic device, the first video data is updated and the second video data is acquired. Obtain the second real-time transmission protocol packet corresponding to the second video data, and send the second real-time transmission protocol packet to the first electronic device.
[0103] According to some embodiments, the media information includes voice information, and the information acquisition unit 604 is used to send the media information to the second electronic device, including: If the format of the voice information does not conform to the encoding negotiated by the communication signaling, the voice information will be decoded into PCM format voice information. The PCM format speech information is re-encoded into a negotiated encoding format speech information; Send voice information in the negotiated encoding format to the second electronic device.
[0104] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0105] Figure 7 This is a schematic diagram of the second type of communication device provided in the embodiments of this disclosure, as shown below. Figure 7 As shown, it includes: The second information receiving unit 701 is used to receive first input information on the dial pad of the first electronic device while the first electronic device is in a call. The collection receiving unit 702 is used to receive and display the collection of text information sent by the network device, and to receive the selection instruction corresponding to the collection of text information. The collection of text information is the collection of information obtained by the network device through recognition processing and conversion processing based on the first input information sent by the first electronic device. The instruction sending unit 703 is used to send a selection instruction to the network device. The selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device. The media information is used to instruct the second electronic device to display the media information. The second electronic device is the peer device in the call process.
[0106] According to some embodiments, the second information receiving unit 701 is further configured to: If no first input information is received on the dial pad of the first electronic device, preset video information is displayed on the screen of the first electronic device.
[0107] According to some embodiments, the second information receiving unit 701 is further configured to: Get the first time point of the first key on the dial pad, where the first key is the start key on the dial pad; Get the second time point of the second button on the dial pad, where the second button is the end button on the dial pad; Based on the first and second time points, obtain the instruction type corresponding to the input instruction; Based on the instruction type, retrieve the input instruction corresponding to the dial pad.
[0108] According to some embodiments, the second information receiving unit 701, when obtaining the instruction type corresponding to the input instruction based on the first time point and the second time point, is specifically used for: Based on the first and second time points, obtain the duration between the first and second time points; If the duration is less than the first duration threshold, the instruction type corresponding to the input instruction is obtained as the input instruction type; If the duration is greater than the second duration threshold, the instruction type corresponding to the input instruction is obtained as a control instruction type, wherein the second duration threshold is greater than or equal to the first duration threshold.
[0109] According to some embodiments, when the collection receiving unit 702 is used to display the collection of text information sent by the network device, it is specifically used for: When displaying a set of text information and receiving a trigger command for a preset key on the dial pad, the set of text information is displayed using a display method corresponding to the preset key.
[0110] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0111] Figure 8This is a schematic diagram of a third type of communication device provided in the embodiments of this disclosure, as shown below. Figure 8 As shown, it includes: The third information receiving unit 801 is used to receive media information sent by the network device. The media information is media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set and the conversion processing of the target text information. Information display unit 802 is used to display media information.
[0112] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0113] This disclosure also provides a network device that can achieve... Figures 1 to 3B The call method described in the document.
[0114] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device 900 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0115] like Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, optical disk, etc.; and a communication unit 909, such as a network card, modem, wireless transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the above methods can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the above methods by any other suitable means (e.g., by means of firmware).
[0116] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0117] Program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0120] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0121] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0122] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0123] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for making a call, characterized in that, Applied to network devices, including: Receive first input information sent by a first electronic device that is in the process of a call, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; The first input information is identified and converted to obtain a set of text information corresponding to the first input information. Send the set of text information to the first electronic device; According to the selection instruction sent by the first electronic device for the set of text information, the target text information is obtained; The target text information is converted to obtain the media information corresponding to the target text information; The media information is sent to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the first electronic device and the second electronic device are devices for conducting a call.
2. The method according to claim 1, characterized in that, The method further includes: When the first electronic device initiates a call, or before the first electronic device initiates a call, the system negotiates with the first electronic device to obtain the display method corresponding to the text information set, wherein the display method is used to indicate the video format for displaying the text information set.
3. The method according to claim 2, characterized in that, The step of obtaining the set of text information corresponding to the first input information includes: The first input information is used to draw an image on the canvas, and the image is decoded to obtain the first video data in the video format; During the process of acquiring the first video data, if no second input information is received from the first electronic device, the first real-time transmission protocol packet corresponding to the first video data is acquired and sent to the first electronic device.
4. The method according to claim 3, characterized in that, The method further includes: During the process of acquiring the first video data, upon receiving the second input information sent by the first electronic device, the first video data is updated to acquire the second video data; Obtain the second real-time transmission protocol packet corresponding to the second video data, and send the second real-time transmission protocol packet to the first electronic device.
5. The method according to claim 1, characterized in that, in, The media information includes voice information, and sending the media information to the second electronic device includes: If the format of the voice information does not conform to the encoding negotiated by the communication signaling, the voice information is decoded into PCM format voice information. The PCM format speech information is re-encoded into a negotiated encoding format speech information; The negotiated encoded voice information is sent to the second electronic device.
6. A method for making a call, characterized in that, Applied to a first electronic device, including: While the first electronic device is in the middle of a call, the first input information is received on the dial pad of the first electronic device; Receive and display a set of text information sent by a network device, and receive a selection instruction corresponding to the set of text information, wherein the set of text information is a set of information obtained by the network device through recognition and conversion processing based on the first input information sent by the first electronic device; The selection instruction is sent to the network device, wherein the selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the second electronic device is the peer device in the call process.
7. The method according to claim 6, characterized in that, The method further includes: If the first input information is not received on the dial pad of the first electronic device, preset video information is displayed on the screen of the first electronic device.
8. The method according to claim 6, characterized in that, The method further includes: Obtain the first time point of the first button on the dial pad, wherein the first button is the start button on the dial pad; Obtain the second time point of the second button on the dial pad, wherein the second button is the end button on the dial pad; Based on the first time point and the second time point, obtain the instruction type corresponding to the input instruction; Based on the instruction type, obtain the input instruction corresponding to the dial pad.
9. The method according to claim 8, characterized in that, The step of obtaining the instruction type corresponding to the input instruction based on the first time point and the second time point includes: Based on the first time point and the second time point, obtain the duration between the first time point and the second time point; If the duration is less than the first duration threshold, the instruction type corresponding to the input instruction is obtained as the input instruction type; If the duration is greater than the second duration threshold, the instruction type corresponding to the input instruction is determined to be a control instruction type, wherein the second duration threshold is greater than or equal to the first duration threshold.
10. The method according to claim 6, characterized in that, The set of text information sent by the network device includes: When displaying the set of text information and receiving a trigger command for a preset key on the dial pad, the set of text information is displayed using a display method corresponding to the preset key.
11. A method for making a call, characterized in that, Applied to a second electronic device, including: The network device receives media information sent by a network device, wherein the media information is media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set, and the target text information is converted and processed. Display the aforementioned media information.
12. A communication device, characterized in that, include: The first information receiving unit is configured to receive first input information sent by a first electronic device during a call, wherein the first input information is the first input information received by the first electronic device on the dial pad of the first electronic device; The set acquisition unit is used to perform recognition and conversion processing on the first input information to acquire the set of text information corresponding to the first input information; A collection sending unit is used to send the collection of text information to the first electronic device; The information acquisition unit is used to acquire target text information according to the selection instruction sent by the first electronic device for the text information set; The information acquisition unit is also used to convert the target text information to obtain the media information corresponding to the target text information; An information sending unit is used to send the media information to a second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the first electronic device and the second electronic device are devices for conducting a call.
13. A communication device, characterized in that, include: The second information receiving unit is used to receive first input information on the dial pad of the first electronic device when the first electronic device is in a call. A collection receiving unit is used to receive and display a collection of text information sent by a network device, and to receive a selection instruction corresponding to the collection of text information, wherein the collection of text information is a collection of information obtained by the network device through recognition processing and conversion processing based on the first input information sent by the first electronic device; The instruction sending unit is used to send the selection instruction to the network device, wherein the selection instruction is used to instruct the network device to obtain target text information according to the selection instruction, perform conversion processing on the target text information, obtain media information corresponding to the target text information, and send the media information to the second electronic device, wherein the media information is used to instruct the second electronic device to display the media information, and the second electronic device is the peer device in the call process.
14. A communication device, characterized in that, include: The third information receiving unit is used to receive media information sent by the network device, wherein the media information is media information corresponding to the target text information obtained by the network device according to the selection instruction sent by the first electronic device for the text information set, and the target text information is converted and processed. An information display unit is used to display the media information.
15. A network device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the call method as described in any one of claims 1 to 5.
16. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the call method as described in any one of claims 6 to 10 or 11.
17. A storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the call method as described in any one of claims 1 to 5, 6 to 10 or 11.