A method for dividing a voice channel of an IP phone

By identifying the media stream direction and timestamp in the IP telephony terminal, and combining it with a large model for channel segmentation and speech transcription, the problem of IP telephony's inability to effectively separate mixed-channel recordings is solved, achieving low-cost, high-efficiency recording data processing and directional analysis.

CN122268984APending Publication Date: 2026-06-23SHANGHAI BINCAI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI BINCAI INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-04-30
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing IP telephony cannot effectively separate mixed-channel recording data and lacks directional analysis capabilities, resulting in high recording transcription costs and making it difficult to meet users' needs for more detailed processing of call content.

Method used

In IP telephony terminals, the direction of media stream flow is identified, and the hardware characteristics of the IP telephony itself are used to divide the audio channels. Combined with timestamp calibration and large model for reasoning and memory, independent channel division and speech transcription of mixed-channel recording data are realized.

Benefits of technology

It reduces the cost of channel segmentation and speech transcription, improves the accuracy of channel separation and the efficiency of speech recognition, supports targeted analysis, and meets users' needs for more detailed processing of call content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122268984A_ABST
    Figure CN122268984A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of IP telephones, in particular to a sound channel division method of an IP telephone, which is applied to an IP telephone terminal and comprises the following steps: recording mixed sound channel recording data generated by interactive multi-parties of the IP telephone in a dialogue interaction process through a recording module of the IP telephone, and uploading the mixed sound channel recording data to a mixed sound channel data separation module; identifying and calibrating the flow direction of a media stream in the received mixed sound channel recording data through the mixed sound channel data separation module of the IP telephone, and dividing the mixed sound channel recording data according to the flow direction to obtain independent sound channel recording data representing a conversation mode, wherein the conversation mode data is used for inputting into a large model and performing inference and memory through the large model. The mixed sound channel recording data is divided into sound channels at the IP telephone terminal, and the calculation cost of subsequent inference and memory of the large model is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of IP telephony technology, and in particular to a method for channel division in IP telephony. Background Technology

[0002] In today's digital communication era, IP telephony (telephones that communicate over a network based on IP addresses), as a communication terminal device, has been widely used due to its advantages such as low cost, high flexibility, and rich functionality. However, existing IP phones can usually only record mixed-channel audio in a simple way, lacking mixed-channel separation and directional analysis capabilities, and thus cannot meet users' needs for more detailed processing of call content.

[0003] Currently, there are technologies on the market that use intelligent algorithms such as large models to separate and transcribe mixed-channel recording data. However, these technologies have high requirements for data processing capabilities and generally require additional hardware and software support, which increases the cost of recording transcription. This creates an irreconcilable contradiction with the original intention of using IP phones (IP telephony) to reduce costs.

[0004] Therefore, there is an urgent need for a low-cost hybrid recording channel separation and transcription system for IP phones to achieve the goal of low-cost, high-performance recording data transcription. Summary of the Invention

[0005] The purpose of this application is to provide a method for channel segmentation in IP telephony, focusing on methods for channel segmentation in IP telephony, aiming to overcome the shortcomings of existing technologies and provide a more efficient and accurate solution for call content processing.

[0006] In some embodiments, this application provides a channel segmentation method for IP telephony, applied in an IP telephony terminal. The method includes: recording mixed-channel audio data generated during dialogue interaction between multiple parties in the IP telephony using the recording module of the IP telephony, and uploading the mixed-channel audio data to a mixed-channel data separation module; identifying and calibrating the flow direction of the media stream in the received mixed-channel audio data using the mixed-channel data separation module of the IP telephony, and segmenting the mixed-channel audio data into channels according to the flow direction to obtain independent channel audio data representing the conversation mode, wherein the conversation mode data is used to input into a large model for inference and memorization.

[0007] In some embodiments, the mixed-channel data separation module further includes a media stream flow direction determination unit; the flow direction determination unit is used to identify and mark the inflow direction and outflow direction of the media stream, and to divide the mixed-channel recording data into channels according to the inflow direction and the outflow direction to obtain independent channel recording data.

[0008] In some embodiments, the mixed channel data separation module further includes a media stream flow direction determination unit; the flow direction determination unit is used to identify and label the near-end data and far-end data in the mixed channel recording data, and to divide the mixed channel recording data into channels according to the near-end data and far-end data to obtain independent channel recording data.

[0009] In some embodiments, the method further includes: timestamping the mixed-channel recording data via IP telephony, and sorting the mixed-channel recording data in time sequence according to the timestamps to obtain mixed-channel recording data with time sequence characteristics generated according to the interaction time; or, timestamping the independent-channel recording data via IP telephony, and sorting the independent-channel recording data in time sequence according to the timestamps to obtain independent-channel recording data representing the session mode with time sequence characteristics generated by different interactive objects according to the interaction time.

[0010] In some embodiments, the method further includes: storing independent channel recording data representing the conversation mode separately according to different channels through the recording storage module of the IP phone.

[0011] In some embodiments, this application also provides an interactive system based on IP telephony channel segmentation, the interactive system including an IP telephony client and a server; the IP telephony client is used to obtain independent channel recording data representing a session mode according to the method in any one of the above embodiments, and upload the independent channel recording data to the server; the large model in the server is used to perform inference and memory based on the received independent channel recording data representing a session mode sent by the IP telephony client.

[0012] In some embodiments, the IP phone terminal and the server are connected in communication. The user of the IP phone terminal can view the recording transcription data associated with the phone terminal identifier on the server and perform targeted analysis based on the recording transcription data.

[0013] This application also provides an IP phone, which includes a recording module, a mixed-channel data separation module, and an independent-channel data storage module. The recording module is used to record mixed-channel recording data generated during the dialogue interaction between the multiple parties in the IP phone. The mixed-channel data separation module is communicatively connected to the recording module and is used to receive the mixed-channel recording data recorded by the recording module, and to divide the mixed-channel recording data into channels to obtain independent-channel recording data representing the conversation mode. The conversation mode data is used to input into a large model for inference and memory. The storage module is communicatively connected to the mixed-channel data separation module and is used to store the independent-channel recording data according to different channels.

[0014] In some embodiments, the IP phone further includes a flow direction determination unit; the flow direction determination unit is used to divide the mixed-channel recording data into channels to obtain independent-channel recording data.

[0015] The above embodiments provide a channel partitioning method for IP telephony. Based on the characteristics of IP telephony, a novel scheme is proposed to partition the mixed-channel recording data recorded by the IP telephony terminal into independent channels according to the direction of the media stream in the mixed-channel recording data. This enables low-cost processing of mixed-channel recording data within the IP telephony and low-cost speech transcription of independent channel recording data within a large model, helping users to further target the application of the recording data. Compared with existing independent channel partitioning techniques that rely on complex data processing such as large models, the media stream partitioning scheme provided in this application not only ensures the accuracy of channel partitioning based on the hardware characteristics of the IP telephony itself, but also greatly reduces the cost of configuring high-end independent channel partitioning hardware and the cost of large model computing power requirements, making its application market prospects very promising. Attached Figure Description

[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:

[0017] Figure 1 This is a flowchart illustrating a method for dividing the audio channels of an IP phone according to one embodiment of this application;

[0018] Figure 2 This is a schematic diagram of a module for an IP phone provided in one embodiment of this application;

[0019] Figure 3 This is a schematic diagram of a module of an interactive system based on IP telephony channel division provided in one embodiment of this application;

[0020] Figure 4 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0022] The technical solutions of the various embodiments of this application can be combined with each other, but only if they are based on the ability of a person skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this application.

[0023] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0024] See Figure 1 This application provides a channel partitioning method for IP telephony, applied in an IP telephony terminal, the method comprising:

[0025] Step S101: Record the mixed-channel audio recording data generated during the dialogue interaction between the multiple parties in the IP phone through the recording module of the IP phone, and upload the mixed-channel audio recording data to the mixed-channel data separation module.

[0026] Step S102: The mixed channel data separation module of the IP phone identifies and calibrates the flow direction of the media stream in the received mixed channel recording data, and divides the mixed channel recording data into channels according to the flow direction to obtain independent channel recording data representing the conversation mode. The conversation mode data is used to input into the large model for inference and memory through the large model.

[0027] In some embodiments, the mixed-channel data recording module of an IP phone terminal typically uses low-cost recording equipment to record the audio data generated during the dialogue interaction between the multiple parties in the IP phone. At this time, the audio channels of the obtained recording data are mixed, and there is no mature channel separation method that can effectively divide the mixed-channel recording data into independent channel data. Instead, some more complex and costly algorithms are used to divide the data into independent channels, which brings a high cost challenge to IP phones.

[0028] In other scenarios, some IP phones have limited recording storage capacity, short storage time, and lack recording-to-text functionality. Developing such functionality separately would be costly. The solution provided in this application operates at the source, recording directly from the IP phone device. It also supports customized channel separation functionality on the IP phone terminal, uploading the independent channel recording data representing the conversation mode to a large model for subsequent data analysis. This reduces the burden of subsequent data analysis on the large model, ensuring low and controllable costs.

[0029] In some embodiments, the mixed-channel data separation module can identify and calibrate the media stream direction of mixed-channel recording data. Specifically, this module can identify the flow direction of the media stream by analyzing the characteristics of the mixed-channel recording data. For example, by detecting changes in the energy and spectral characteristics of the audio signal, it can determine whether the audio is emitted locally or received remotely. Simultaneously, the source and destination address information in the IP protocol can be combined to further determine the direction of the media stream. After identifying the direction of the media stream, it is calibrated using specific markers for subsequent channel segmentation operations. Specifically, the channel segmentation method of the mixed-channel data separation module can divide the mixed-channel recording data into independent channel recording data based on the flow direction of the media stream. In other embodiments, digital signal processing techniques, such as adaptive filtering algorithms, can be further employed to process the mixed-channel audio signal and separate the audio signals from different directions. For example, the minimum mean square error (LMS) adaptive filtering algorithm can be used to continuously adjust the filter coefficients so that the output signal approximates the audio signals of the independent channels as closely as possible.

[0030] The mixed-channel data separation module is used to identify and calibrate the flow direction of the media stream in the received mixed-channel recording data, and divide the mixed-channel recording data into channels according to the flow direction to obtain independent channel recording data, and upload the independent channel recording data to the large model. Further, the mixed-channel data separation module is also used to timestamp the independent channel recording data, and sort the independent channel recording data in time sequence according to the timestamps to obtain independent channel recording data representing the conversation mode, and upload the independent channel recording data representing the conversation mode to the large model.

[0031] In some embodiments, the mixed-channel data separation module has a timestamp calibration function.

[0032] In some embodiments, the independent channel recording data is timestamped via IP telephony, and then sorted chronologically according to the timestamps to obtain independent channel recording data representing conversation patterns with temporal characteristics generated by different interactive objects according to their interaction time. Specifically, this module adds a precise timestamp to the audio segment following each independent channel recording data. The timestamp can be obtained based on a system clock or network clock to ensure time accuracy. For example, a high-precision atomic clock or network time protocol is used to synchronize time, ensuring that the timestamp error is within milliseconds. Further chronological sorting is performed, and the independent channel recording data is sorted chronologically according to the timestamps to obtain independent channel recording data representing conversation patterns.

[0033] In some embodiments, a sorting algorithm, such as quicksort or mergesort, can be used to sort the timestamps of the independent channel recording data to ensure the efficiency and accuracy of the sorting.

[0034] In another embodiment, this application may further include timestamping the mixed-channel recording data via IP telephony, and sorting the mixed-channel recording data chronologically according to the timestamps to obtain mixed-channel recording data with temporal characteristics generated according to the interaction time. Specifically, after the IP telephony recording module of this application records the mixed-channel recording data, it can first timestamp the mixed-channel recording data, and then further divide the time-stamped mixed-channel recording data into independent channels through a mixed-channel data separation module, thus obtaining independent-channel recording data representing the conversation mode.

[0035] It can be understood that independent channel recording data representing a conversation pattern refers to recording data that is distinguished according to the participants and the conversation time. In other words, the mixed channel data obtained from the original IP telephony recording is divided into independent channels according to the participants and timestamped according to the conversation time of the participants, so as to obtain independent channel recording data that can distinguish both the participants and the conversation time, representing the conversation pattern.

[0036] In some embodiments, after the obtained independent channel recording data representing the conversation pattern is transmitted to a large model, the large model can use advanced speech recognition techniques to perform speech recognition on the sorted recording data. For example, deep learning models, such as Long Short-Term Memory Networks (LSTM) or Gated Recurrent Units (GRUs) based on Recurrent Neural Networks (RNNs), can be used to extract features from the audio signals and perform speech recognition. Simultaneously, language models, such as statistical language models or neural network language models, can be combined to improve the accuracy of speech recognition. After recognizing the speech content, it is transcribed into text data for convenient subsequent analysis and processing.

[0037] In modern communication networks, IP telephony has become an important means of communication for businesses and individuals. Effective processing of its audio data is crucial for applications such as call record analysis and customer service quality monitoring. However, existing telephone recording and speech recognition systems have significant limitations. Traditional IP telephony systems often simply record a mixture of the voices of both parties, failing to separate the voices of different speakers. This makes it difficult to accurately distinguish between different speakers when analyzing the call content, posing significant challenges to information extraction and analysis. For example, in customer service scenarios, the inability to clearly distinguish between the voices of customer service personnel and customers makes it difficult to accurately assess the service quality of customer service personnel and the needs of customers. Moreover, the lack of targeted analysis capabilities prevents the processing of call content according to specific needs, failing to meet users' demands for more detailed processing of call content.

[0038] This application addresses the shortcomings of existing technologies, which primarily focus on simply recording mixed-channel audio without separating the mixed channels or performing speech recognition and transcription. It creatively proposes a convenient and effective technique that enables low-cost, pre-emptive separation of mixed-channel recording data into independent channel recording data representing the conversation pattern at the IP telephony end. Since time stamping and channel segmentation are performed at the IP telephony end, it's understood that the IP telephony end, being closest to both parties and the first to acquire the original interaction data, makes channel segmentation and timestamping at this end the simplest, most convenient, and least data-loss-prone method. The method provided in this application fully utilizes the inherent hardware advantages of IP telephony during channel segmentation, achieving accurate and effective channel segmentation without incurring excessive additional costs. Compared to channel segmentation through complex data processing algorithms at the large-scale model end, this application segments channels based on the media stream flow direction at the IP end, significantly improving the convenience, low cost, and effectiveness of channel segmentation. Compared to embedding costly hardware with channel segmentation capabilities into IP phones, this application utilizes a simple and effective concept of media flow direction for channel segmentation. This not only controls the cost advantage of IP phones but also allows current IP phones to fully leverage their existing capabilities and empower new functionalities. This is undoubtedly a highly competitive product highlight for IP phone terminals.

[0039] Furthermore, this application achieves independent channel segmentation of mixed-channel recording data at the IP telephony end with almost no additional hardware cost. The data transmitted to the large model is thus independent channel recording data representing the conversation pattern. The large model no longer needs to deploy a separate channel segmentation algorithm, reducing the computational demands on the large model and eliminating the need to optimize the algorithm within the large model to improve the accuracy of independent channel segmentation. Directional analysis is also supported within the large model, significantly improving the efficiency and accuracy of call content processing. Moreover, this application addresses the limitation of existing IP telephony technologies that only provide a vague, mixed audio stream, while the IP telephony provided by this application can clearly present the content and chronological order of each speaker's remarks, offering users more valuable information.

[0040] The above embodiments provide a channel segmentation method for IP telephony. Based on the characteristics of IP telephony recording data, a novel scheme is proposed to segment the mixed-channel recording data recorded by the IP telephony terminal into independent channels according to the direction of the media stream in the mixed-channel recording data. This enables speech transcription of the mixed-channel recording data, helping users to further target the application of the recording data. Compared with existing independent channel segmentation techniques that rely on complex data processing such as large models, this media stream segmentation scheme not only ensures the accuracy of channel segmentation but also significantly reduces the cost of independent channel segmentation, making its application market prospects very promising.

[0041] In some specific embodiments, the mixed channel data separation module further includes a media stream flow direction determination unit; the flow direction determination unit is used to identify and mark the inflow and outflow directions of the media stream, and to divide the mixed channel recording data into channels according to the inflow and outflow directions to obtain independent channel recording data.

[0042] In the above embodiment, a mixed-channel recording separation is achieved by determining the direction of the media stream. The media stream flow direction determination unit in the mixed-channel data separation module of this IP phone employs a more precise method to identify and label the incoming and outgoing directions of the media stream. Specifically, the flow direction determination unit has the function of identifying the incoming and outgoing directions of the media stream, which can be comprehensively determined by combining the characteristics of the mixed-channel recording data and network communication information. For example, by analyzing the energy distribution and spectral characteristics of the mixed-channel recording data, it can be determined whether the mixed-channel recording data is sent locally or received remotely. Simultaneously, the transmission direction of the media stream is determined using the source and destination address information in the IP protocol. If the source address is a local IP address and the destination address is a remote IP address, the media stream can be determined to be outgoing; otherwise, it is incoming. The channel division function of the flow direction determination unit can specifically divide the mixed-channel recording data into channels based on the identified incoming and outgoing directions of the media stream. For example, the audio signals in the incoming and outgoing directions can be allocated to different channels. For example, the incoming audio signal is assigned to the left channel, and the outgoing audio signal is assigned to the right channel, thus obtaining independent channel recording data. Specifically, the mixed channel recording data can be divided into channels according to the flow direction, and the different channel data obtained from the division can be stored in the corresponding independent channels to obtain independent channel recording data.

[0043] Existing technologies often employ complex techniques to divide mixed-channel data, which requires significant data processing resources and incurs high costs. This application, based on the characteristics of IP telephony, creatively combines the media stream direction in mixed-channel recording data and divides the channels according to the inflow and outflow directions of the media stream. This approach can more accurately and cost-effectively distinguish audio data from different directions, thereby improving the accuracy of channel separation.

[0044] The above embodiments have at least the following beneficial effects. First, they improve the accuracy of channel separation. Based on the two-way communication characteristics of IP telephony, the incoming and outgoing directions of the media stream can be accurately identified, enabling more accurate division of mixed-channel recording data into independent channel recording data, reducing interference between channels, and improving the quality of channel separation. Second, they enhance speech recognition and transcription effects. Accurate channel separation provides clearer audio data for subsequent large-scale model speech recognition and transcription, helping to improve the accuracy of speech recognition and the quality of transcription, thereby better presenting the call content and conducting subsequent targeted analysis.

[0045] In some embodiments, the mixed channel data separation module further includes a media stream flow direction determination unit, which is used to determine the interaction device of the interaction object, and to calibrate the flow direction of the media stream in the mixed channel recording data according to the interaction device of different interaction objects, and to take the media streams generated by different interaction devices as different flow directions of the media stream.

[0046] The above embodiment specifically describes an IP telephony system that calibrates the direction of media streams based on interactive devices. The media stream flow direction determination unit in the mixed-channel data separation module of this system can calibrate the flow direction of the media stream in the mixed-channel recording data according to the interactive devices of different interactive objects. In this embodiment, the flow direction determination unit has the function of identifying and determining interactive devices. In specific implementations, the interactive device of the interactive object can be determined through device identification information, device type information, and network connection information. For example, in an IP telephony system, each device has a unique device identifier (such as a MAC address), and the type and affiliation of the device can be determined by querying the device identifier database. Simultaneously, combined with network connection information, such as IP address and port number, the communication relationship between devices can be further determined. The flow direction determination unit has the function of calibrating the direction of media streams, treating the media streams generated by different interactive devices as different flow directions of the media streams. For example, if two interactive devices A and B are in a call, the media stream corresponding to the audio signal emitted by device A is in one direction, and the media stream corresponding to the audio signal emitted by device B is in another direction. By calibrating the media streams of different interactive devices, the audio data of different interactive objects can be distinguished more accurately. The flow direction determination unit has a channel division function, which can divide mixed-channel recording data into channels according to the calibrated media flow direction. For example, a channel allocation method based on device identification can be used to assign audio signals from different interactive devices to different channels. For example, the audio signal of device A can be assigned to the left channel, and the audio signal of device B can be assigned to the right channel, thus obtaining independent channel recording data.

[0047] The above embodiments, based on existing technologies, do not consider the influence of interactive devices on the direction of media streams, and therefore cannot distinguish audio data based on different interactive devices. The technical solution of this application, however, combines the characteristics of IP telephony to determine the direction of media streams based on the interactive devices, enabling more accurate differentiation of audio data from different interactive objects and improving the channel separation effect.

[0048] The technical solutions provided in the above embodiments of this application have at least the following beneficial effects: First, they can accurately distinguish the audio of different interactive objects. Based on the characteristics of IP telephony itself, by calibrating the media stream direction according to the interactive device, the audio data of different interactive objects can be accurately separated, making it convenient for users to analyze and process the speech content of different speakers separately. Second, they can improve the adaptability of channel separation. In complex call scenarios, multiple interactive devices may participate in the call simultaneously. This solution can accurately distinguish the media stream direction according to the different interactive devices, improving the adaptability and accuracy of channel separation in complex scenarios.

[0049] In some embodiments, the mixed channel data separation module further includes a media stream flow direction determination unit; the flow direction determination unit is used to identify and label the near-end data and far-end data in the mixed channel recording data, and to divide the mixed channel recording data into channels according to the near-end data and far-end data to obtain independent channel recording data.

[0050] The above embodiment specifically describes an IP phone that divides audio channels based on near-end and far-end data. The media stream flow direction determination unit in the mixed-channel data separation module of this IP phone can identify and label near-end and far-end data in the mixed-channel recording data. Specifically, in this embodiment, the flow direction determination unit has the function of identifying near-end and far-end data. For example, it can identify near-end and far-end data through the intensity, spectral characteristics, and echo information of the audio signal. For example, near-end data usually has higher audio intensity and richer low-frequency components, while far-end data has relatively lower audio intensity and may have echo phenomena. By analyzing these characteristics, it can be determined whether the audio data comes from a near-end device or a far-end device. The flow direction determination unit also has a channel division function, which can divide the mixed-channel recording data into channels based on near-end and far-end data. For example, adaptive filtering and signal separation algorithms can be used to extract near-end and far-end data separately and allocate them to different channels. For example, near-end data can be allocated to the left channel, and far-end data to the right channel, thereby obtaining independent channel recording data.

[0051] The above embodiments, based on existing technology, do not distinguish between near-end and far-end data in IP telephony, and cannot perform channel segmentation based on audio information from different locations. The technical solution of this application utilizes the characteristics of IP telephony to identify near-end and far-end data for channel segmentation, which can better restore audio information from different locations and improve the accuracy of channel separation. In this embodiment, the location characteristics of audio data are further considered, making channel separation more consistent with reality.

[0052] The embodiments provided in this application have at least the following beneficial effects: First, they can better restore audio information. By distinguishing between near-end and far-end data for channel segmentation, they can more accurately restore audio information from different locations, allowing users to more clearly understand the voice characteristics and positional relationship between the two parties in the call. Second, they can improve the accuracy of speech recognition and transcription. Clear channel separation provides cleaner audio data for speech recognition and transcription, which helps to improve the accuracy of speech recognition, thereby more accurately transcribing the call content.

[0053] In some embodiments, the mixed-channel data recording module is installed in the IP phone and records the interaction data of the multiple parties in the IP phone during the dialogue interaction process to form mixed-channel recording data.

[0054] The above embodiments specifically provide an IP phone with an integrated recording module. This system integrates a mixed-channel data recording module within the IP phone, employing an integrated design that combines audio acquisition, encoding, and storage functions. In a specific embodiment, the mixed-channel data recording module has an audio acquisition function, capable of capturing audio through a built-in microphone to ensure clear and accurate audio quality. The mixed-channel data recording module also has an encoding function, capable of encoding the acquired audio data and converting it into digital audio signals. Furthermore, the mixed-channel data recording module has a data storage function, capable of storing the encoded audio data in the IP phone's internal storage device, such as flash memory or a hard drive. Simultaneously, storage capacity thresholds and data cleanup strategies can be set to ensure efficient use of storage space.

[0055] In the above embodiments, considering that existing technologies typically separate the recording module from the IP phone, requiring additional equipment and connections to implement the recording function, the technical solution provided in this application integrates the recording module into the IP phone, reducing the number and complexity of devices and improving the system's integration and ease of use. Furthermore, the mixed-channel data recording module provided in this application is the simplest recording module, capable only of generating mixed-channel recording data, and lacks mature technology for direct channel classification; therefore, the cost of this recording function is very low.

[0056] The technical solution provided in the above embodiments of this application has at least the following beneficial effects: First, in terms of data acquisition, users do not need to connect additional recording equipment; they can complete the recording operation simply by using an IP phone, greatly improving the convenience of data acquisition. Second, it improves system integration; the integrated design tightly integrates the recording module with the IP phone, reducing communication interfaces and connection lines between devices, and improving system stability and reliability. Third, it enhances ease of use; users can directly set and manage recordings on the IP phone, making operation simpler and more convenient, and lowering the user's barrier to entry. More importantly, for low-cost IP phones, the configured mixed-channel data recording module is relatively inexpensive and cannot achieve the division of independent channels, which brings significant cost and technical challenges to subsequent recording data transcription. The technical solution provided in this application is well-suited to the low-cost scenario of IP phones, creatively proposing a channel division technology based on media stream direction recognition in IP phone terminals by combining the device characteristics and interactive features of IP phones. This achieves low-cost and efficient division and recognition of mixed-channel data recordings, providing the necessary information source for subsequent speech transcription and directional analysis.

[0057] In some embodiments, the large model includes a recording recognition and transcription unit; the recording recognition and transcription unit is used to perform speech recognition on the sorted independent channel recording data to obtain speech transcription data with temporal characteristics generated by different interactive objects according to the interaction time.

[0058] The above embodiments provide an IP telephony capable of data interaction with a large model. The recording data transcription module provided by this large model employs advanced speech recognition and natural language processing (NLP) technologies in its recording recognition and transcription unit. Specifically, the recording data transcription module has speech recognition capabilities. For example, it can use deep learning models, such as a hybrid model based on convolutional neural networks (CNN) and long short-term memory networks (LSTM), to perform speech recognition on sorted, independent channel recording data. This model can automatically extract features from the audio signal and convert them into text information. Simultaneously, it can be trained using large-scale speech datasets to improve the model's recognition accuracy and generalization ability. The recording data transcription module also possesses NLP capabilities. Building upon speech recognition, it can use NLP techniques to process the recognized text, improving its readability and accuracy. For example, it can use techniques such as part-of-speech tagging, syntactic analysis, and semantic understanding to perform grammatical checking and semantic analysis on the text, correcting recognition errors and supplementing missing information. Furthermore, the audio transcription module features a temporal feature presentation function. By associating and sorting the recognized speech transcription data with timestamps, it presents speech transcription data with temporal features generated by different interactive objects according to the interaction time. Visualization tools can be used to display the speech transcription data in the form of a timeline, making it convenient for users to view and analyze.

[0059] The embodiments provided in this application can effectively present temporal feature data and perform in-depth processing of recognition results. Through a dedicated recording recognition and transcription unit, and by employing advanced speech recognition and natural language processing technologies, the efficiency and accuracy of speech recognition and transcription are improved, and the temporal features of the dialogue can be better presented.

[0060] The above embodiments have at least the following beneficial effects: First, they can improve the efficiency of speech recognition and transcription. By employing advanced deep learning models and natural language processing technology, audio data can be quickly and accurately converted into text information, thus improving the efficiency of speech recognition and transcription. Second, they can enhance the accuracy of speech recognition and transcription. Through large-scale dataset training and the application of natural language processing technology, recognition errors can be effectively corrected, improving the accuracy of speech recognition and transcription. Third, they can clearly present the temporal characteristics of the dialogue. By presenting the speech transcription data in the form of a timeline, users can clearly understand the speaking time and order of different interactive parties, and better grasp the overall process of the dialogue.

[0061] In some embodiments, see Figure 3This application also provides an interactive system based on IP telephony channel segmentation, the interactive system including an IP telephony terminal and a server terminal; the IP telephony terminal is used to obtain independent channel recording data representing the session mode according to the method provided in any of the above embodiments, and upload the independent channel recording data to the server terminal; the large model in the server terminal is used to perform inference and memory based on the received independent channel recording data representing the session mode sent by the IP telephony terminal.

[0062] In some embodiments, the IP phone terminal and the server are connected in communication. The user of the IP phone terminal can view the recording transcription data associated with the phone terminal identifier on the server and perform targeted analysis based on the recording transcription data.

[0063] In some embodiments, the IP phone terminal communicates with the server, and the user of the IP phone terminal can view the recording data associated with the phone terminal identifier on the server and perform targeted analysis based on the recording data.

[0064] The above embodiments provide an IP telephony interaction system that supports targeted analysis. Specifically, the IP telephony terminal and the server communicate via a secure and reliable network connection. IP telephony terminal users can obtain recording data associated with a specific terminal identifier from the server and perform targeted analysis on the recording data according to specific analysis rules. For example, each IP telephony terminal has a unique and private user ID. By logging in with the user ID, all recording data under that ID can be found on the server, allowing for targeted analysis. This targeted analysis can include early warning analysis, such as providing advance warnings of potential dangers during calls, and business opportunity discovery, such as analyzing call data to predict user needs and make targeted recommendations.

[0065] In the above embodiments, users of IP phone terminals can log in to their IP phone accounts to search for all recorded data associated with their IP phone on the server. Specifically, they can send a request to the server carrying the phone terminal's identification information. The server then queries its database based on this identification information, finds the associated recorded data, and returns it to the terminal for viewing. To ensure data security and privacy, data encryption and authentication technologies can be used to protect the data transmission process. Furthermore, the server can also have targeted analysis capabilities, allowing for the setting of targeted analysis rules based on different application scenarios and user needs. For example, in customer service scenarios, rules can be set to analyze customer service personnel's service attitude, response time, and problem resolution rate; in business negotiation scenarios, rules can analyze the tone, stance, and key demands of both parties. For example, data mining and machine learning algorithms can be used to analyze the recorded data. For instance, sentiment analysis algorithms can be used to analyze the speaker's emotional tendencies, keyword extraction algorithms can be used to extract key information from the dialogue, and clustering analysis algorithms can be used to classify and summarize the dialogue.

[0066] Specifically, users interact via IP phones, recording the conversation and creating a mixed-channel recording file. The IP phone then divides the recording file into independent channels and sends it to a speech recognition server via an interface. A large-scale speech recognition model is used to obtain a chronologically ordered voice chat log. The server grants users account access to their IP phones. Enterprise users configure their IP phone accounts with usernames and passwords, allowing them to store the recording data on the server. Users can then manage their recordings, view all recordings under their account, and analyze the data.

[0067] In the above embodiments, the IP phone is connected to the server for communication. By utilizing the server's powerful data analysis capabilities, in-depth data analysis functions are provided to IP phone users. This supports the separation of mixed-channel recording data, transcription of recording data, and targeted analysis, providing IP users with significant commercial value.

[0068] It should be noted that the IP phone client and server in this application can be two independent ends. The server is an independent data analysis platform that can connect to a large number of IP phone clients. Utilizing the server's powerful data analysis and processing capabilities, it can perform separate and centralized data analysis on the recording data from each IP phone client, and further synchronize the processed data to the IP phone users, allowing them to conduct data mining and analysis. In other words, this application provides a server platform for centrally processing IP phone recording data, which can effectively, conveniently, and cost-effectively process and analyze mixed IP phone recording data.

[0069] In some embodiments, see Figure 2 This application also provides an IP phone, which includes a recording module, a mixed-channel data separation module, and an independent-channel data storage module. The recording module is used to record mixed-channel recording data generated during the dialogue interaction between the multiple parties in the IP phone. The mixed-channel data separation module is communicatively connected to the recording module and is used to receive the mixed-channel recording data recorded by the recording module, and to divide the mixed-channel recording data into channels to obtain independent-channel recording data representing the conversation mode. The conversation mode data is used to input into a large model for inference and memory. The storage module is communicatively connected to the mixed-channel data separation module and is used to store the independent-channel recording data according to different channels.

[0070] In some embodiments, the IP phone further includes a flow direction determination unit; the flow direction determination unit is used to divide the mixed-channel recording data into channels to obtain independent-channel recording data.

[0071] The specific functions of each module in the IP telephony terminal are not described in detail; for specific implementation examples of each module, please refer to the descriptions of other embodiments in the text.

[0072] It is understood that the computer device for the channel partitioning method for IP telephony provided in this application can be a server, and its internal structure diagram can be as follows: Figure 3 As shown. The computer device includes a processor, memory, and network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores relevant data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the method provided in this application. It should be noted that the recording data involved in this application is all user-authorized data.

[0073] Those skilled in the art will understand that Figure 4The structures shown are merely block diagrams of a portion of the structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. The computer device may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program may include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0074] It should be understood that the processor mentioned in the embodiments of this application can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0075] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0076] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for channel partitioning in IP telephony, characterized in that, The method, applied to IP telephony terminals, includes: The recording module of the IP phone records the mixed-channel audio data generated during the dialogue interaction between the multiple parties in the IP phone, and uploads the mixed-channel audio data to the mixed-channel data separation module. The mixed-channel data separation module of the IP phone identifies and calibrates the flow direction of the media stream in the received mixed-channel recording data, and divides the mixed-channel recording data into channels according to the flow direction to obtain independent channel recording data representing the conversation mode. The conversation mode data is used to input into the large model for inference and memory.

2. The method according to claim 1, characterized in that, The hybrid audio channel data separation module also includes a media stream flow direction determination unit; The flow direction determination unit is used to identify and mark the inflow and outflow directions of the media stream, and to divide the mixed-channel recording data into channels according to the inflow and outflow directions to obtain independent channel recording data.

3. The method according to any one of claims 1 or 2, characterized in that, The mixed-channel data separation module also includes a media stream flow direction determination unit; The flow direction determination unit is used to identify and calibrate the near-end data and far-end data in the mixed-channel recording data, and to divide the mixed-channel recording data into channels according to the near-end data and far-end data to obtain independent channel recording data.

4. The method according to claim 1, characterized in that, The method further includes: The mixed-channel recording data is timestamped using IP telephony, and then sorted chronologically according to the timestamps to obtain mixed-channel recording data with temporal characteristics generated based on interaction time; or, The independent channel recording data is timestamped using IP telephony, and then sorted chronologically according to the timestamps to obtain independent channel recording data representing conversation patterns with temporal characteristics generated by different interactive objects according to their interaction time.

5. The method according to claim 1, characterized in that, The method further includes: The recording storage module of the IP phone stores the independent channel recording data that represents the conversation mode separately according to different channels.

6. The method according to claim 1, characterized in that, The method further includes: The mixed-channel recording data is divided into channels according to the flow direction, and the different channel data obtained are stored in the corresponding independent channels to obtain independent channel recording data.

7. An interactive system based on IP telephony channel segmentation, characterized in that, The interactive system includes an IP telephony client and a server. The IP phone terminal is used to obtain independent channel recording data representing the session mode by the method according to any one of claims 1 to 6, and uploads the independent channel recording data to the server. The large model on the server side is used for inference and memory based on the independent channel recording data representing the session pattern sent by the IP phone.

8. The interactive system according to claim 7, characterized in that, The IP phone terminal communicates with the server terminal. The user of the IP phone terminal can view the recording transcription data associated with the phone terminal identifier on the server terminal and perform targeted analysis based on the recording transcription data.

9. An IP telephony, characterized in that, The IP phone includes a recording module, a mixed-channel data separation module, and an independent-channel data storage module; The recording module is used to record mixed-channel audio data generated during the dialogue interaction between multiple parties in an IP phone call. The mixed-channel data separation module is communicatively connected to the recording module and is used to receive the mixed-channel recording data recorded by the recording module, and to divide the mixed-channel recording data into channels to obtain independent channel recording data representing the conversation mode. The conversation mode data is used to input into the large model for inference and memory through the large model. The storage module is communicatively connected to the mixed-channel data separation module and is used to store the independent channel recording data separately according to different channels.

10. The IP telephony according to claim 9, characterized in that, The IP phone also includes a flow direction determination unit; The flow direction determination unit is used to divide the mixed channel recording data into channels to obtain independent channel recording data.