Network call recording method, device and equipment and readable storage medium

By mixing and encoding the raw audio stream of a VoIP call in the terminal device to generate a network call recording file, the problems of poor universality and high cost in the existing technology are solved, and a cross-system recording solution is realized.

CN121644726APending Publication Date: 2026-03-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies have poor versatility and high costs in VoIP call recording, and cannot be applied to various call systems.

Method used

By acquiring and mixing the raw audio streams of the first and second terminals in the terminal device, a network call recording file is generated, avoiding the dependence on dedicated recording equipment.

Benefits of technology

It improves the versatility of VoIP call recording, reduces recording costs, and is no longer limited to specific call systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644726A_ABST
    Figure CN121644726A_ABST
Patent Text Reader

Abstract

The invention discloses a network call recording method, device and equipment and a readable storage medium, and the method comprises the steps: obtaining a first original audio stream when a first terminal and a second terminal establish a network call connection; the first original audio stream is an audio stream generated by the first terminal based on the first object voice in the network call connection process; the first terminal receives a transmission audio stream sent by the second terminal through the network call connection, and performs audio transmission decoding processing on the transmission audio stream to obtain a second original audio stream; the second original audio stream is an audio stream generated by the second terminal based on the second object voice in the network call connection process; and carrying out audio mixed coding processing on the first original audio stream and the second original audio stream to obtain a call recording file aiming at the network call connection. According to the invention, the versatility of the network call recording can be improved, and the cost of the network call recording can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and readable storage medium for recording network calls. Background Technology

[0002] VoIP, or Internet Protocol (IP) calling, is a voice communication technology that uses the Internet Protocol (IP) to enable voice calls and multimedia conferencing—in other words, communication via the internet. With the increasing penetration and widespread use of the internet, more and more users prefer to conduct VoIP calls directly through communication applications, leading to a growing demand for VoIP call recording. For example, corporate call centers can record VoIP calls to assess the quality of their external consultations and services, which is highly beneficial for improving their service capabilities.

[0003] Current technology for recording VoIP calls typically involves adding a relay, specifically an additional RTP (Real-time Transport Protocol) relay within the voice communication network. This allows the RTP stream, which meets recording requirements, to be redirected to dedicated call recording equipment for capture before reaching the other end of the call. However, this approach is only applicable to specific call systems, has extremely poor versatility, and is very costly to implement. Summary of the Invention

[0004] This application provides a method, apparatus, device, and readable storage medium for recording network calls, which can improve the versatility of network call recording and reduce the cost of network call recording.

[0005] This application provides a method for recording network calls, which is executed by a first terminal and includes:

[0006] When the first terminal and the second terminal establish a network call connection through a communication application, the first original audio stream is obtained; the network call connection is established based on the first communication application in the first terminal and the second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the voice of the first object of the first object during the network call connection process; the first object refers to the object that logs in to the first communication application in the first terminal through the first object account;

[0007] The system receives the transmitted audio stream sent by the second terminal through a network call connection, performs audio transmission decoding processing on the transmitted audio stream, and obtains the second original audio stream. The second original audio stream refers to the audio stream generated by the second terminal based on the voice of the second object during the network call connection. The second object refers to the object that logs into the second communication application in the second terminal through the second object's account.

[0008] The first and second original audio streams are subjected to audio mixing and encoding to obtain a call recording file for the network call connection.

[0009] One embodiment of this application provides a network call recording device, which is operated by a first terminal and includes:

[0010] The acquisition module is used to acquire a first raw audio stream when the first terminal establishes a network call connection with the second terminal; the network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal; the first raw audio stream refers to the audio stream generated by the first terminal based on the voice of a first object during the network call connection process; the first object refers to the object that logs into the first communication application in the first terminal through the first object account;

[0011] The decoding module is used to receive the transmitted audio stream sent by the second terminal through a network call connection, perform audio transmission decoding processing on the transmitted audio stream, and obtain the second original audio stream; the second original audio stream refers to the audio stream generated by the second terminal based on the voice of the second object during the network call connection; the second object refers to the object that logs into the second communication application in the second terminal through the second object account;

[0012] The hybrid encoding module is used to perform audio hybrid encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection.

[0013] In one possible implementation, the acquisition module is also used to perform the following operations:

[0014] Retrieve the audio notification playback file; the audio notification playback file is an audio file generated for the audio notification playback.

[0015] A mixed audio stream is generated based on the first original audio stream and the recording notification broadcast file;

[0016] The mixed audio stream is sent to the second terminal so that the second terminal plays the recording notification voice while playing the voice of the first object, based on the mixed audio stream.

[0017] When a recording confirmation notification for the recording notification voice message is received from the second terminal, the step of performing audio mixing and encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection is executed.

[0018] In one possible implementation, when the acquisition module generates the mixed transmission audio stream based on the first raw audio stream and the recording notification playback file, it specifically performs the following operations:

[0019] Generate the original recording notification playback file based on the recording notification playback file;

[0020] A recording notification broadcast audio stream is generated based on the recording notification broadcast file; the frame size of the audio frames contained in the recording notification broadcast audio stream is the same as the frame size of the audio frames contained in the first original audio stream.

[0021] The first original audio stream and the recording notification broadcast audio stream are mixed to obtain a mixed audio stream.

[0022] The mixed audio stream is processed by audio transmission encoding to obtain the mixed transmission audio stream.

[0023] In one possible implementation, when the acquisition module generates the original recording notification playback file based on the recording notification playback file, it specifically performs the following operations:

[0024] Perform format recognition processing on the recorded notification playback file to obtain the format recognition result;

[0025] If the format recognition result is a lossless compressed audio file format, then the recording notification playback file will be identified as the original recording notification playback file.

[0026] If the format recognition result is a lossy compressed audio file format, then the recording notification broadcast file is subjected to audio decoding processing to obtain the original recording notification broadcast file.

[0027] In one possible implementation, when the acquisition module generates the audio stream for recording notification playback based on the recording notification playback file, it specifically performs the following operations:

[0028] Get the audio frame size of the first raw audio stream;

[0029] The file data contained in the recording notification playback file is read and processed according to the audio frame size to obtain N read file data; the data size of each read file data is equal to the audio frame size;

[0030] Based on the audio frame format information of the first original audio stream, the formats of the N read file data are converted respectively to obtain N recording notification audio frames;

[0031] N audio frames of the recording notification broadcast are identified as the audio stream of the recording notification broadcast.

[0032] In one possible implementation, the acquisition module is used to read and process the file data contained in the recording notification playback file according to the audio frame size. When N pieces of file data are obtained, it is specifically used to perform the following operations:

[0033] Generate a read buffer based on the audio frame size;

[0034] The file pointer is moved 1st time to start reading the recording notification playback file from the position it points to in the i-th move, and the file data read in the i-th move is obtained; the file pointer is moved 1st time to point to the starting position of the recording notification playback file.

[0035] If the amount of file data read in the i-th reading is greater than the size of an audio frame, then the file data read in the i-th reading is written sequentially into the read buffer, the data contained in the read buffer is determined as the i-th file data to be read, and the read buffer is cleared.

[0036] Based on the audio frame size, move the file pointer i-th time backward in the recording notification playback file to obtain the file pointer i+1-th time backward, and continue to read the recording notification playback file from the position pointed to by the file pointer i+1-th time backward.

[0037] If the amount of file data read in the i-th time is less than or equal to the size of the audio frame, then the file data read in the i-th time is written sequentially into the read buffer, the remaining buffer area of ​​the read buffer is filled with default data, the data contained in the read buffer is determined as the i-th read file data, and the first i read file data is determined as N read file data.

[0038] In one possible implementation, the acquisition module is used to mix the first original audio stream and the recording notification playback audio stream to obtain the mixed audio stream, specifically for performing the following operations:

[0039] The first original audio stream and the recording notification broadcast audio stream are added together to obtain a summed audio stream;

[0040] The summed audio stream is normalized to obtain the mixed audio stream.

[0041] In one possible implementation, the mixing encoding module performs audio mixing encoding on the first and second original audio streams to obtain a call recording file for the network call connection, specifically performing the following operations:

[0042] The first original audio stream is converted to the left channel to obtain the left channel audio stream;

[0043] The second original audio stream is converted to the right channel to obtain the right channel audio stream.

[0044] The left and right channel audio streams are merged to obtain a stereo audio stream.

[0045] Audio encoding processing is performed on the two-channel audio stream to obtain a call recording file for the network call connection.

[0046] In one possible implementation, the mixing encoding module is used to merge the left and right channel audio streams to obtain a stereo audio stream. Specifically, it performs the following operations:

[0047] Iterate through the left channel audio stream to obtain the j-th left channel audio frame and the j-th right channel audio frame in the right channel audio stream; the j-th left channel audio frame and the j-th right channel audio frame are time-aligned; j is a positive integer;

[0048] The target left channel audio frame and the target right channel audio frame are alternately merged to obtain the j-th stereo audio frame containing 2M sample points; M is a positive integer.

[0049] When the traversal of the left channel audio frames contained in the left channel audio stream is complete, a two-channel audio stream is generated based on the obtained two-channel audio frames.

[0050] In one possible implementation, the hybrid encoding module is also used to perform the following operations:

[0051] When a network call connection is detected to be disconnected, the call information corresponding to the network call connection is retrieved;

[0052] A call archive file is generated based on call information and call recording files, and then sent to an archive server for storage. The archive server, upon receiving a file retrieval request containing filtered call information from a quality inspection server, sends the target call archive file to the quality inspection server. The target call archive file refers to the call archive file containing call information that matches the filtered call information. Upon receiving the target call archive file, the quality inspection server performs quality inspection analysis on it to obtain the analysis results for the filtered call information.

[0053] In one possible implementation, when the acquisition module acquires the first raw audio stream, it specifically performs the following operations:

[0054] During the network call connection, the microphone component is used to collect the voice of the first party to obtain an analog voice electrical signal;

[0055] The analog voice signal is converted from analog to digital to obtain the first original audio stream.

[0056] In one possible implementation, the hybrid encoding module is also used to perform the following operations:

[0057] Obtain the call information corresponding to the network call connection; the call information includes the account information of the first object account, the account information of the second object account, and the network call connection environment information;

[0058] The call information is processed by feature extraction to obtain the first object account features, the second object account features, and the network call connection environment features;

[0059] The current call scenario is obtained by performing call scenario prediction processing based on the characteristics of the first object account, the characteristics of the second object account, and the characteristics of the network call connection environment.

[0060] If the current call scenario is a recorded call scenario, then the step of performing audio mixing and encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection is executed.

[0061] One embodiment of this application provides a computer device, including: a processor, a memory, and a network interface;

[0062] The processor is connected to the memory and the network interface. The network interface is used to provide a data communication network element, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method in the embodiments of this application.

[0063] One aspect of this application provides a computer-readable storage medium storing a computer program adapted for loading by a processor and executing the methods described in this application.

[0064] One aspect of this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in this application.

[0065] In this embodiment, when a first terminal establishes a network call connection with a second terminal, the first terminal can acquire a first original audio stream, then receive a transmitted audio stream sent by the second terminal through the network call connection, perform audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream, and finally perform audio mixing encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection. Here, the network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the first object's voice during the network call connection process; the first object refers to the object that logs into the first communication application in the first terminal through the first object's account; the second original audio stream refers to the audio stream generated by the second terminal based on the second object's voice during the network call connection process; the second object refers to the object that logs into the second communication application in the second terminal through the second object's account. The method provided in this embodiment uses a terminal to complete the recording of network calls, eliminating the need for dedicated call recording equipment, greatly saving recording costs, and is no longer limited to a specific call system, significantly improving the versatility of network call recording. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application;

[0068] Figure 2 This is a schematic diagram illustrating an application scenario of a network call recording method provided in an embodiment of this application;

[0069] Figure 3 This is a flowchart illustrating a network call recording method provided in an embodiment of this application;

[0070] Figure 4 This is a schematic diagram illustrating a scenario for generating a call recording file according to an embodiment of this application;

[0071] Figure 5 This is a flowchart illustrating a network call recording method provided in an embodiment of this application;

[0072] Figure 6 This is a schematic diagram of a scenario for adding voice broadcast according to an embodiment of this application;

[0073] Figure 7 This is a schematic diagram of the structure of a network call recording device provided in an embodiment of this application;

[0074] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0076] To facilitate understanding, the following brief explanations are provided for some of the terms:

[0077] 1. PCM (Pulse Code Modulation) is a standard method for representing analog signals as digital signals. It is one of the most basic and common audio data formats, widely used in various audio processing and transmission systems, such as CDs, telephones, and computer audio. PCM data is typically stored linearly, meaning the amplitude value of each sample point is directly stored as a binary number.

[0078] 2. Sampling Rate: The sampling rate refers to the number of times an audio signal is sampled per second, measured in Hertz (Hz). Common sampling rates include 44.1kHz (CD quality), 48kHz (professional audio and video), and 96kHz. For example, a sampling rate of 44.1kHz means that 44,100 samples are collected per second.

[0079] 3. Frame Size: Frame size refers to the number of samples contained in a single processing or transmission. Audio processing typically divides consecutive samples into frames, with each frame containing a number of samples. The choice of frame size can affect processing latency and efficiency. Smaller frame sizes can reduce latency but may increase processing overhead; larger frame sizes can improve processing efficiency but increase latency.

[0080] 4. Samples per Frame: Samples per frame refers to the number of sample points contained in each frame. This value is usually a function of the frame size and the sampling rate. Assuming the sampling rate is R (unit: Hz), the duration of each frame is T (unit: seconds), and the frame size is N (unit: number of samples), it can be expressed as: N = T x R.

[0081] 5. Opus is a lossy audio encoding format that aims to contain sound and speech in a single format and is suitable for low-latency, real-time audio transmission over the network.

[0082] Please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application. Figure 1 As shown, this network architecture may include a communication server 2000 and a terminal cluster. The terminal cluster may specifically include one or more terminals; the number of terminals in the terminal cluster is not limited here. Figure 1 As shown, the multiple terminals may specifically include terminal 3000a, terminal 3000b, terminal 3000c, ..., terminal 3000n; terminal 3000a, terminal 3000b, terminal 3000c, ..., terminal 3000n can be directly or indirectly connected to server 2000 via wired or wireless communication, so that each terminal can interact with server 2000 through the network connection.

[0083] Each terminal in the terminal cluster can include: smartphones, tablets, laptops, desktop computers, smart voice interaction devices, smart home appliances (e.g., smart TVs), wearable devices, in-vehicle terminals, aircraft, and other smart terminals with data processing capabilities. It should be understood that, as... Figure 1 Each terminal in the terminal cluster shown can have a communication application client installed. When this application client runs on each terminal, it can communicate with the aforementioned... Figure 1 The servers 2000 shown interact with each other. The communication application refers to third-party platform-integrated applications (mini-programs) and mobile applications (APPs) that provide network calling functionality. This application client can be a standalone client or an embedded sub-client integrated into another client (such as an instant messaging client, social networking client, video client, etc.); this is not limited here. Different terminals can have the same communication application installed, or they can have different communication applications installed.

[0084] Among them, such as Figure 1The server 2000 shown can be the server corresponding to the application client. This server 2000 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. For example, when... Figure 1 When each terminal in the terminal cluster shown has the same communication application installed, such as Figure 1 The server 2000 shown can be the backend application server corresponding to this communication application; when Figure 1 When each terminal in the terminal cluster shown has a different communication application installed, such as Figure 1 The server 2000 shown can contain a backend server for each communication application.

[0085] Among them, such as Figure 1 Each terminal in the terminal cluster shown can transmit data with server 2000 through an installed communication application, thereby establishing a communication connection with other terminals for network calls. Taking the establishment of a communication connection between terminals 3000a and 3000b as an example, terminal 3000a can generate a call request through an installed first communication application and then transmit the call request to server 2000. Server 2000 can then forward the call request to terminal 3000b. When server 2000 receives a call consent request from terminal 3000b generated through an installed second communication application, server 2000 can simultaneously establish data transmission connections with both terminals 3000a and 3000b. The data transmission connections between server 2000 and terminal 3000a, and between server 2000 and terminal 3000b, constitute the communication connection between terminals 3000a and 3000b. The first and second communication applications can be the same or different applications.

[0086] Furthermore, embodiments of this application may be implemented in... Figure 1 Among the multiple terminals shown, one terminal is selected as the first terminal, and another terminal is selected as the second terminal. For example, in the embodiments of this application, one terminal can be selected as the second terminal. Figure 1Terminal 3000a is shown as the first terminal, and terminal 3000b is shown as the second terminal. The first terminal may have a first communication application installed, and the second terminal may have a second communication application installed. Both the first and second terminals can interact with the server 2000 through the first and second communication applications. Furthermore, in this embodiment, the user corresponding to the terminal can be referred to as an object. For example, the user corresponding to the first terminal can be referred to as the first object, and the user corresponding to the second terminal can be referred to as the second object.

[0087] like Figure 1 As shown, when a network call is conducted between a first object and a second object, and the first object needs to record the network call, the network call recording method provided in this application can be used to record the network call through the first terminal. Specifically, the network call recording method is as follows: when the first terminal establishes a network call connection with the second terminal, the first terminal can obtain a first original audio stream. The network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal. The first original audio stream refers to the audio stream generated by the first terminal based on the first object's voice during the network call connection process. The first object refers to the object that logs into the first communication application in the first terminal through the first object's account. The first object's voice refers to the voice generated by the first object during the call. Simultaneously, the first terminal can receive the transmitted audio stream sent by the second terminal through the network call connection, and perform audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream. The second original audio stream refers to the audio stream generated by the second terminal based on the second object's voice during the network call connection process. The second object refers to the object that logs into the second communication application in the second terminal through the second object's account. The second object's voice refers to the voice generated by the second object during the call. Finally, the first terminal can perform audio mixing and encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection.

[0088] It is understood that, in the specific embodiments of this application, the data related to audio streams, etc., when applied to specific products or technologies, requires user permission or consent, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.

[0089] It is understood that the embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, blockchain, and intelligent driving.

[0090] Furthermore, for a better understanding of the implementation process of the above-mentioned network call recording, please refer to [link to relevant documentation]. Figure 2This is a schematic diagram illustrating an application scenario of a network call recording method provided in an embodiment of this application. Wherein, as... Figure 2 The backend server 20 shown can be the one described above. Figure 1 The server 2000 in the corresponding embodiment, such as Figure 2 The terminal devices 200 and 201 shown can be the aforementioned Figure 1 Any two different terminal devices in the terminal device cluster in the corresponding embodiment, for example, terminal device 200 can be terminal device 3000a and terminal device 201 can be terminal device 3000b.

[0091] like Figure 2 As shown, taking an enterprise service scenario as an example, in order to provide better enterprise services, enterprises usually have dedicated customer service representatives to solve problems for different customers. With the widespread use of VoIP, more and more customers prefer to conduct VoIP calls directly with enterprise customer service through communication applications. Enterprises usually need to record and archive each VoIP call in order to check the quality of the consultation and services provided. Assume that terminal devices 200 and 201 both have the same communication application installed. The user associated with terminal device 200 is customer A, and the user associated with terminal device 201 is customer service representative B of enterprise C. In order to improve the quality of enterprise C's subsequent enterprise services, when customer A and customer service representative B conduct a VoIP call, customer service representative B can notify customer A that the entire VoIP call will be recorded, and then record the VoIP call with customer A's consent.

[0092] like Figure 2 As shown, after terminal device 200 and terminal device 201 establish a network call connection through a communication application, data interaction can be performed through this network call connection to realize a network call between customer A and customer service B. Specifically, terminal device 201 can generate a raw audio stream 2011 based on customer service B's voice during the call connection, then perform audio transmission encoding processing on it to obtain a transmission audio stream 2012 for transmission, and then send the transmission audio stream 2012 to terminal device 200 through the network call connection. Terminal device 200 can play customer service B's voice by decoding the received transmission audio stream 2012. Here, the raw audio stream refers to a lossless encoded digital signal used to represent analog signals, such as PCM data. The transmission audio stream refers to a lossy encoded digital signal used for real-time audio transmission, such as OPUS data. Similarly, terminal device 201 can also receive the transmission audio stream 2001 sent by terminal device 200 through the network call connection, and play customer A's voice by decoding the transmission audio stream 2001.

[0093] like Figure 2As shown, when terminal device 201 determines that it needs to record the current network call, it can acquire the original audio stream 2011 generated during the network call connection and simultaneously acquire the transmitted audio stream 2001 received through the network call connection. It then performs audio transmission decoding processing on the transmitted audio stream 2001 to obtain the original audio stream 2002. This original audio stream 2002 can be generated by terminal device 200 based on customer A's voice during the call connection. Then, terminal device 201 can perform audio mixing encoding processing on the original audio stream 2001 and the original audio stream 2002 to obtain a call recording file 2013 for the network call connection, specifically the call recording file 2013 for the network call between customer A and customer service representative B.

[0094] like Figure 2 As shown, after detecting a disconnection in the network call, the terminal device 201 can upload the generated final call recording file 2013 to a dedicated storage server 21 for storage. It can be understood that the storage server 21 is mainly used to store call recording files generated by terminals corresponding to multiple customer service representatives, and to organize and store them. Furthermore, the enterprise can also use the enterprise customer service quality inspection system in the enterprise server 22 to periodically perform quality inspections on the multiple call recording files stored in the storage server 21. That is, by analyzing the call recording files, potential service problems can be identified, allowing the enterprise to optimize service processes in real time and improve customer satisfaction.

[0095] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a network call recording method provided in an embodiment of this application. The method can be implemented by a first terminal (e.g., the one described above). Figure 1 The method is executed by terminal 3000a) in the corresponding embodiment. The following description will take the execution of this method by the first terminal as an example, wherein the network call recording method may include at least the following steps S101-S103:

[0096] Step S101: When the first terminal establishes a network call connection with the second terminal, the first original audio stream is obtained; the network call connection is established based on the first communication application in the first terminal and the second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the first object's voice during the network call connection process; the first object refers to the object that logs into the first communication application in the first terminal through the first object account.

[0097] Specifically, the first party can initiate a call request to the second party through a first communication application logged in on a first terminal, or the second party can initiate a call request to the first party through a second communication application logged in on a second terminal. When the second party accepts the call request from the first party through the second communication application logged in on the second terminal, or when the first party accepts the call request from the second party through the first communication application logged in on the first terminal, a network call connection is established between the first terminal and the second terminal to enable a network call between the first party and the second party. The network call connection refers to the data transmission channel (or data transmission connection) used to transmit the audio stream generated during the network call. When the first communication application and the second communication application are the same communication application, they correspond to the same backend server. In this case, the network call connection can include a data transmission channel between the first terminal and the backend server, and a data transmission channel between the backend server and the second terminal. The audio stream generated by the first terminal can be transmitted to the backend server via the data transmission channel between the first terminal and the backend server, and then transmitted to the second terminal via the data transmission channel between the backend server and the second terminal. Similarly, the audio stream generated by the second terminal can be transmitted to the backend server via the data transmission channel between the backend server and the second terminal, and then transmitted to the first terminal via the data transmission channel between the first terminal and the backend server. First terminal; when the first communication application and the second communication application are different communication applications, they correspond to different backend servers. In this case, the network call connection can include a data transmission channel between the first terminal and the first backend server corresponding to the first communication application, a data transmission channel between the second terminal and the second backend server corresponding to the second communication application, and a data transmission channel between the first backend server and the second backend server. At this time, the audio stream generated by the first terminal can be transmitted to the first backend server through the data transmission channel between the first terminal and the first backend server, then transmitted to the second backend server through the data transmission channel between the first backend server and the second backend server, and then transmitted to the second terminal through the data transmission communication between the second backend server and the second terminal. Similarly, the audio stream generated by the second terminal can be transmitted to the first terminal.

[0098] Specifically, the first raw audio stream refers to the raw audio stream generated by the first terminal. The raw audio stream is a lossless encoded digital signal used to represent analog speech signals, such as PCM data. A feasible implementation of the first terminal acquiring the first raw audio stream can be as follows: During a network call connection, the first terminal acquires the speech of the first recipient through a microphone component to obtain an analog speech signal; the analog speech signal undergoes analog-to-digital conversion (ADC) processing to obtain the first raw audio stream. Here, the analog speech signal is an electrical signal converted from a sound signal, used to imitate and analogize the characteristics of sound in the real world. However, similar to sound, electrical signals cannot be directly stored, as electricity requires a continuous energy source. Therefore, the analog speech signal needs to undergo ADC processing to convert it into a storable digital signal. This ADC processing can be implemented using an ADC (Analog-to-Digital Converter), which is a device that converts continuously changing analog signals into discrete digital signals. It is understandable that when the terminal needs to play the first object's voice, the first original audio stream can be processed by digital-to-analog conversion, that is, the stored digital signal can be converted into an analog voice signal through a DAC (Digital to analog converter), and then the analog voice signal can be used to drive the speaker to play the first object's voice.

[0099] Step S102: Receive the transmitted audio stream sent by the second terminal through the network call connection, and perform audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream; the second original audio stream refers to the audio stream generated by the second terminal based on the second object's voice during the network call connection process; the second object refers to the object that logs into the second communication application in the second terminal through the second object's account.

[0100] Specifically, lossy audio formats are typically used in real-time audio transmission, such as Opus, MP3 (Moving Picture Experts Group Audio Layer III), and AAC (Advanced Audio Coding). Therefore, before transmission, the original audio stream usually needs to undergo audio transmission encoding processing, converting it from a lossless digital signal to a lossy digital signal. Thus, the transmitted audio stream received by the first terminal is obtained after the second terminal has processed the second original audio stream using audio transmission encoding. The first terminal needs to perform audio transmission decoding processing on the transmitted audio stream to obtain the second original audio stream.

[0101] Step S103: Perform audio mixing and encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

[0102] Specifically, in order to improve the clarity of the call recording file during subsequent playback, one of the original audio streams can be used as the left channel and the other as the right channel. The audio streams are mixed and encoded in real time. In order to save storage space, a lossy encoding format (mp3 or aac, etc.) is generally used. That is, the final call recording file can be an mp3 file or an aac file.

[0103] Specifically, a feasible implementation process for performing audio mixing and encoding on the first and second original audio streams to obtain a call recording file for a network call connection can be as follows: performing left channel conversion on the first original audio stream to obtain a left channel audio stream; performing right channel conversion on the second original audio stream to obtain a right channel audio stream; performing audio stream merging on the left and right channel audio streams to obtain a stereo audio stream; and performing audio encoding on the stereo audio streams to obtain a call recording file for a network call connection.

[0104] Specifically, a feasible implementation process for merging the left and right channel audio streams to obtain a stereo audio stream can be as follows: Iterate through and obtain the j-th left channel audio frame from the left channel audio stream and the j-th right channel audio frame from the right channel audio stream; the j-th left channel audio frame and the j-th right channel audio frame are time-aligned; j is a positive integer; alternately merge the M sample points contained in the target left channel audio frame and the M sample points contained in the target right channel audio frame to obtain the j-th stereo audio frame containing 2M sample points; M is a positive integer; when the traversal of the left channel audio frames contained in the left channel audio stream is complete, generate a stereo audio stream based on the obtained stereo audio frames. The stereo audio stream contains all the obtained stereo audio frames. For example, the left channel audio frame Z contains two sampling points: [Z1, Z2], and the right channel audio frame Y, which is time-aligned with it, contains two sampling points: [Y1, Y2]. The resulting stereo audio frame ZY, obtained by alternately merging the sampling points of the two frames, contains [Z1, Y1, Z2, Y2]. It can be understood that in the above process, the left and right channel audio streams are alternately stored to generate a stereo audio stream. This alternating storage method is based on how the human auditory system processes stereo audio, that is, by judging the direction and distance of the sound source through the phase difference of the sound received by the left and right ears, and thus forming a stereo audio perception in the brain. This ensures that the information of the left and right channels is correctly processed and restored during storage and transmission, thereby maintaining the stereo effect of the final stereo audio stream.

[0105] To better understand the process of generating the call recording files, please refer to the example of an enterprise service scenario. Figure 4 , Figure 4 This is a schematic diagram illustrating a scenario for generating a call recording file according to an embodiment of this application. For example... Figure 4 As shown, after terminal 41 collects customer voice 401, it generates customer transmitted audio stream 402 based on it. Then, customer transmitted audio stream 402 is transmitted over the network and sent to terminal 42. Terminal 42 can perform audio transmission decoding on the received customer transmitted audio stream 402 to obtain the customer original audio stream 403, and then perform right channel conversion on it to obtain the right channel audio stream 404.

[0106] like Figure 4 As shown, while receiving the customer's transmitted audio stream 402, terminal 42 also collects and converts the customer service voice 405 through the installed microphone component, and then obtains the original customer service audio stream 406. Terminal 42 performs left channel conversion on the original customer service audio stream 406 to obtain the left channel audio stream 407. Then, terminal 42 performs audio stream encoding on the left channel audio stream 407 and the right channel audio stream 404 together to obtain the final call recording file 408.

[0107] Optionally, when a network call connection is detected to be disconnected, the call information corresponding to the network call connection is obtained; a call archive file is generated based on the call information and the call recording file, and the call archive file is sent to the archive server for storage. The call information may include call start time, call duration, call participant information, call terminal information, etc. The archive server, upon receiving a file retrieval request containing filtered call information from the quality inspection server, sends the target call archive file to the quality inspection server; the target call archive file refers to the call archive file containing call information that matches the filtered call information; the quality inspection server, upon receiving the target call archive file, performs quality inspection analysis on the target call archive file to obtain the quality inspection analysis results for the filtered call information. In other words, the quality inspection server can periodically perform spot checks and analyses on the call storage files stored on the storage server. For example, in an enterprise service scenario, the enterprise needs to periodically check the quality of the consultations and services provided to external parties. Suppose the enterprise needs to check the service performance of Customer Service 1, it can send a quality inspection request for Customer Service 1 to the quality inspection server. The quality inspection server can use Customer Service 1's object information as the filter call information, and then send a file retrieval request containing the filter call information to the archive server. After receiving the file retrieval request, the quality inspection server can use all call archive files containing Customer Service 1's object information as target call archive files, and then send the target call archive files to the quality inspection server. The quality inspection server can perform quality inspection analysis on the target call archive files, such as call duration analysis, call satisfaction analysis, etc., to obtain the quality inspection analysis results for Customer Service 1 and determine the service performance of Customer Service 1.

[0108] Optionally, obtain the call information corresponding to the network call connection; the call information includes the account information of the first object account, the account information of the second object account, and the network call connection environment information; perform feature extraction processing on the call information to obtain the features of the first object account, the features of the second object account, and the features of the network call connection environment; perform call scenario prediction processing based on the features of the first object account, the features of the second object account, and the features of the network call connection environment to obtain the current call scenario; if the current call scenario belongs to a recording call scenario, then execute step S103. The call scenario prediction processing can be implemented using a classification model, that is, a call scenario prediction model for scenario prediction is pre-trained, and then the extracted features are input into the call scenario prediction model. Based on the output call scenario prediction label, it can be determined whether the current call scenario belongs to a recording call scenario.

[0109] The method provided in this application allows for recording of network calls, whether based on the same or different communication applications, through a terminal. This eliminates the need to deploy dedicated recording equipment for specific calling systems, significantly reducing recording costs and greatly improving the versatility of network call recording. Furthermore, since audio stream encoding and decoding occur at the edge terminal without server involvement, server resource consumption is reduced, further lowering operating costs. Moreover, distributing encoding and decoding tasks across various edge terminals makes the system more flexible and scalable; even if one terminal malfunctions, it will not affect the operation of the entire system.

[0110] Further, please see Figure 5 , Figure 5 This is a flowchart illustrating a network call recording method provided in an embodiment of this application. The method can be implemented by a first terminal (e.g., the one described above). Figure 1 The method is executed by terminal 3000a) in the corresponding embodiment. The following description will take the execution of this method by the first terminal as an example, wherein the network call recording method may include at least the following steps S201-S206:

[0111] Step S201: When the first terminal establishes a network call connection with the second terminal, a first original audio stream is obtained; the network call connection is established based on the first communication application in the first terminal and the second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the first object's voice during the network call connection process; the first object refers to the object that logs into the first communication application in the first terminal through the first object account.

[0112] Step S202: The first terminal receives the transmitted audio stream sent by the second terminal through the network call connection, performs audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream; the second original audio stream refers to the audio stream generated by the second terminal based on the second object's voice during the network call connection; the second object refers to the object that logs into the second communication application in the second terminal through the second object's account.

[0113] Specifically, the implementation of steps S201-S202 can be found above. Figure 3 The specific descriptions of steps S101-S102 in the corresponding embodiments will not be repeated here.

[0114] Step S203: Obtain the recording notification playback file; the recording notification playback file is an audio file generated for the recording notification playback voice.

[0115] Specifically, the recorded notification file can be a pre-recorded file stored directly on the first terminal, containing digital signals generated based on the recorded notification voice. The recorded notification voice is used to remind the user that the current network call will be recorded; for example, the recorded notification voice could be, "Dear user, in order to provide better service in the future, this call will be recorded."

[0116] Step S204: Generate a mixed transmission audio stream based on the first original audio stream and the recording notification broadcast file.

[0117] Specifically, a feasible implementation process for generating a mixed transmission audio stream based on the first original audio stream and the recording notification broadcast file can be as follows: generating an original recording notification broadcast file based on the recording notification broadcast file; generating a recording notification broadcast audio stream based on the recording notification broadcast file; the frame size of the audio frames contained in the recording notification broadcast audio stream is the same as the frame size of the audio frames contained in the first original audio stream; performing mixing processing on the first original audio stream and the recording notification broadcast audio stream to obtain a mixed audio stream; and performing audio transmission encoding processing on the mixed audio stream to obtain a mixed transmission audio stream.

[0118] Specifically, a feasible implementation process for generating the original recorded notification broadcast file from the recorded notification broadcast file can be as follows: The recorded notification broadcast file undergoes format recognition processing to obtain a format recognition result; if the format recognition result is a lossless compressed audio file format, then the recorded notification broadcast file is determined as the original recorded notification broadcast file; if the format recognition result is a lossy compressed audio file format, then the recorded notification broadcast file undergoes audio decoding processing to obtain the original recorded notification broadcast file. The lossless compressed audio file format can be PCM format, and the lossy compressed audio file format can be MP3 format, AAC format, etc. That is to say, in the subsequent mixing process, a lossless compressed audio file format digital signal is required. Therefore, if the stored recorded notification broadcast file is a lossy compressed audio file format, it needs to be converted to a lossless compressed audio file format.

[0119] Specifically, not just any two audio streams can be directly mixed. Only two audio streams with the same frame size and audio frame format can be mixed. Therefore, the generated recording notification audio stream must contain audio frames with the same frame size as the first original audio stream, and their audio frame formats must be identical. The audio frame format includes the sampling rate, frame length, bit depth, and number of channels. Therefore, a feasible implementation of generating the recording notification audio stream from the recording notification file can be as follows: obtain the audio frame size of the first original audio stream; read the file data contained in the recording notification file according to the audio frame size to obtain N read file data; the data size of each read file data is equal to the audio frame size; convert the format of each of the N read file data according to the audio frame format information of the first original audio stream to obtain N recording notification audio frames; and determine the N recording notification audio frames as the recording notification audio stream. The audio frame format information may include the sampling rate, number of channels, bit depth, etc.

[0120] Specifically, a feasible implementation process for reading and processing the file data contained in the recording notification playback file according to the audio frame size to obtain N read file data can be as follows: A read buffer is generated based on the audio frame size; the recording notification playback file is read starting from the position pointed to by the i-th file pointer movement to obtain the i-th read file data; the first file pointer movement points to the beginning position of the recording notification playback file; if the amount of file data read in the i-th move is greater than the audio frame size, the file data read in the i-th move is sequentially written into the read buffer, and the data contained in the read buffer is determined as the i-th read file. The process involves clearing the read buffer, moving the file pointer in the recording notification file according to the audio frame size (i-th move) to obtain the (i+1)-th moved file pointer, and continuing to read the recording notification file from the position pointed to by the (i+1)-th moved file pointer. If the amount of data read in the i-th read is less than or equal to the audio frame size, the data read in the i-th read is sequentially written into the read buffer. The remaining buffer area of ​​the read buffer is filled with default data. The data contained in the read buffer is determined as the i-th read file data, and the first i read file data are determined as N read file data. Generating the read buffer according to the audio frame size means creating a read buffer with the size of the audio frame. For example, if each audio frame in the original audio stream contains 320 sample points, and each sample point is 16 bits (2 bytes), then the buffer size is 640 bytes. The recording notification file can be opened and read using file I / O functions, such as the fopen function in C. The default data value is 0. When the end of the file is reached, the number of bytes read may be less than the buffer size. In this case, the remaining part of the buffer can be filled with zeros to ensure the integrity of the audio frames.

[0121] Understandably, in the feasible implementation process of generating N read file data as described above, the file data read in the i-th read will be sequentially written into the read buffer. When the amount of file data read in the i-th read is greater than the audio frame size, only the data written into the read buffer will be determined as the i-th read file data. When the amount of file data read in the i-th read is less than the audio frame size, the read buffer will first undergo default data filling processing, and then the data contained in the read buffer will be determined as the i-th read file data. In this way, it can be ensured that no matter how much or how little file data is read each time, the final determined size of the read file data is always the audio frame size, thereby ensuring the accuracy of the mixed audio stream obtained from subsequent mixing processing. In addition, after each read file data is determined, the file pointer is moved forward by the audio frame size before starting a new round of reading, which can avoid the problem of duplicate data or discontinuous data in the N read file data.

[0122] Specifically, mixing algorithms such as simple additive mixing and weighted mixing can be selected during audio mixing, and this application does not impose any restrictions. For example, when using a simple additive mixing algorithm, a feasible implementation process for mixing the first original audio stream and the recording notification broadcast audio stream to obtain the mixed audio stream can be as follows: add the first original audio stream and the recording notification broadcast audio stream to obtain a summed audio stream; normalize the summed audio stream to obtain the mixed audio stream.

[0123] Step S205: The mixed audio stream is sent to the second terminal so that the second terminal plays the recording notification voice while playing the voice of the first object according to the mixed audio stream.

[0124] Specifically, the second terminal can perform audio transmission decoding processing on the mixed audio stream to obtain the mixed audio stream, and then perform digital-to-analog conversion processing on it to obtain an analog mixed electrical signal. The second terminal can drive the speaker through the analog mixed electrical signal to play the recording notification voice while playing the voice of the first object.

[0125] Step S206: When a confirmation notification for the recording notification broadcast voice sent by the second terminal is received, the first original audio stream and the second original audio stream are subjected to audio mixing encoding processing to obtain a call recording file for the network call connection.

[0126] To facilitate understanding of the process of adding the above-mentioned audio notification playback file, please refer to the example of an enterprise service scenario. Figure 6 , Figure 6 This is a schematic diagram illustrating a scenario for adding voice broadcasting, as provided in an embodiment of this application. Figure 6As shown, after the customer service terminal collects the customer service voice through the microphone, it obtains the customer service audio stream (i.e., the original customer service audio stream mentioned above). Simultaneously, the customer service terminal can obtain the voice broadcast file (i.e., the recorded notification broadcast file mentioned above). Assuming the voice broadcast file is in MP3 format, the customer service terminal can perform audio decoding on the voice broadcast file, that is, convert the voice broadcast file from MP3 format to PCM format. Then, the customer service terminal can loop through the PCM format voice broadcast file to obtain the broadcast audio stream. The specific process of looping through the file can be found in the detailed description of step S204 above, and will not be repeated here. Then, the customer service terminal can mix the customer service audio stream and the broadcast audio stream to obtain a mixed audio stream, and then encode it to obtain the mixed transmission audio stream for transmission. This mixed transmission audio stream is then transmitted to the customer terminal via the network (i.e., the network call connection mentioned above).

[0127] The method provided in this application systematically reminds users that their online calls will be recorded, and the encoding of the call recording file only begins after the user's consent is confirmed, thus ensuring the compliance of the generated call recording file. Furthermore, the audio stream encoding and decoding are performed on the terminal, not the server, which increases system security, reduces the potential attack surface, and makes end-to-end encryption easier to achieve, significantly reducing the risk of data leakage and tampering.

[0128] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a network call recording device provided in an embodiment of this application. The network call recording device can be a computer program (including program code) running on a computer device; for example, the network call recording device is an application software. The network call recording device 1 can be used to execute corresponding steps in the data processing method provided in the embodiment of this application. Figure 7 As shown, the network call recording device 1 may include: an acquisition module 110, a decoding module 120, and a hybrid encoding module 130.

[0129] The acquisition module 110 is used to acquire a first original audio stream when the first terminal establishes a network call connection with the second terminal; the network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the voice of a first object during the network call connection process; the first object refers to the object that logs into the first communication application in the first terminal through the first object account;

[0130] The decoding module 120 is used to receive the transmitted audio stream sent by the second terminal through a network call connection, and to perform audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream; the second original audio stream refers to the audio stream generated by the second terminal based on the voice of the second object during the network call connection; the second object refers to the object that logs into the second communication application in the second terminal through the second object account;

[0131] The hybrid encoding module 130 is used to perform audio hybrid encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

[0132] In one possible implementation, the acquisition module 110 is also used to perform the following operations:

[0133] Retrieve the audio notification playback file; the audio notification playback file is an audio file generated for the audio notification playback.

[0134] A mixed audio stream is generated based on the first original audio stream and the recording notification broadcast file;

[0135] The mixed audio stream is sent to the second terminal so that the second terminal plays the recording notification voice while playing the voice of the first object, based on the mixed audio stream.

[0136] When a recording confirmation notification for the recording notification voice message is received from the second terminal, the step of performing audio mixing and encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection is executed.

[0137] In one possible implementation, when the acquisition module 110 generates the mixed transmission audio stream based on the first original audio stream and the recording notification playback file, it is specifically used to perform the following operations:

[0138] Generate the original recording notification playback file based on the recording notification playback file;

[0139] A recording notification broadcast audio stream is generated based on the recording notification broadcast file; the frame size of the audio frames contained in the recording notification broadcast audio stream is the same as the frame size of the audio frames contained in the first original audio stream.

[0140] The first original audio stream and the recording notification broadcast audio stream are mixed to obtain a mixed audio stream.

[0141] The mixed audio stream is processed by audio transmission encoding to obtain the mixed transmission audio stream.

[0142] In one possible implementation, when the acquisition module 110 generates the original recording notification playback file based on the recording notification playback file, it specifically performs the following operations:

[0143] Perform format recognition processing on the recorded notification playback file to obtain the format recognition result;

[0144] If the format recognition result is a lossless compressed audio file format, then the recording notification playback file will be identified as the original recording notification playback file.

[0145] If the format recognition result is a lossy compressed audio file format, then the recording notification broadcast file is subjected to audio decoding processing to obtain the original recording notification broadcast file.

[0146] In one possible implementation, when the acquisition module 110 generates the audio stream for recording notification playback based on the recording notification playback file, it specifically performs the following operations:

[0147] Get the audio frame size of the first raw audio stream;

[0148] The file data contained in the recording notification playback file is read and processed according to the audio frame size to obtain N read file data; the data size of each read file data is equal to the audio frame size;

[0149] Based on the audio frame format information of the first original audio stream, the formats of the N read file data are converted respectively to obtain N recording notification audio frames;

[0150] N audio frames of the recording notification broadcast are identified as the audio stream of the recording notification broadcast.

[0151] In one possible implementation, the acquisition module 110 is used to read and process the file data contained in the recording notification playback file according to the audio frame size. When N pieces of read file data are obtained, it is specifically used to perform the following operations:

[0152] Generate a read buffer based on the audio frame size;

[0153] The file pointer is moved 1st time to start reading the recording notification playback file from the position it points to in the i-th move, and the file data read in the i-th move is obtained; the file pointer is moved 1st time to point to the starting position of the recording notification playback file.

[0154] If the amount of file data read in the i-th reading is greater than the size of an audio frame, then the file data read in the i-th reading is written sequentially into the read buffer, the data contained in the read buffer is determined as the i-th file data to be read, and the read buffer is cleared.

[0155] Based on the audio frame size, move the file pointer i-th time backward in the recording notification playback file to obtain the file pointer i+1-th time backward, and continue to read the recording notification playback file from the position pointed to by the file pointer i+1-th time backward.

[0156] If the amount of file data read in the i-th time is less than or equal to the size of the audio frame, then the file data read in the i-th time is written sequentially into the read buffer, the remaining buffer area of ​​the read buffer is filled with default data, the data contained in the read buffer is determined as the i-th read file data, and the first i read file data is determined as N read file data.

[0157] In one possible implementation, the acquisition module 110 is used to perform audio mixing processing on the first original audio stream and the recording notification broadcast audio stream, and when obtaining the mixed audio stream, it is specifically used to perform the following operations:

[0158] The first original audio stream and the recording notification broadcast audio stream are added together to obtain a summed audio stream;

[0159] The summed audio stream is normalized to obtain the mixed audio stream.

[0160] In one possible implementation, when the mixing encoding module 130 performs audio mixing encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection, it is specifically used to perform the following operations:

[0161] The first original audio stream is converted to the left channel to obtain the left channel audio stream;

[0162] The second original audio stream is converted to the right channel to obtain the right channel audio stream.

[0163] The left and right channel audio streams are merged to obtain a stereo audio stream.

[0164] Audio encoding processing is performed on the two-channel audio stream to obtain a call recording file for the network call connection.

[0165] In one possible implementation, the mixing encoding module 130 is used to perform audio stream merging processing on the left channel audio stream and the right channel audio stream to obtain a stereo audio stream, specifically for performing the following operations:

[0166] Iterate through the left channel audio stream to obtain the j-th left channel audio frame and the j-th right channel audio frame in the right channel audio stream; the j-th left channel audio frame and the j-th right channel audio frame are time-aligned; j is a positive integer;

[0167] The target left channel audio frame and the target right channel audio frame are alternately merged to obtain the j-th stereo audio frame containing 2M sample points; M is a positive integer.

[0168] When the traversal of the left channel audio frames contained in the left channel audio stream is complete, a two-channel audio stream is generated based on the obtained two-channel audio frames.

[0169] In one possible implementation, the hybrid encoding module 130 is also used to perform the following operations:

[0170] When a network call connection is detected to be disconnected, the call information corresponding to the network call connection is retrieved;

[0171] A call archive file is generated based on call information and call recording files, and then sent to an archive server for storage. The archive server, upon receiving a file retrieval request containing filtered call information from a quality inspection server, sends the target call archive file to the quality inspection server. The target call archive file refers to the call archive file containing call information that matches the filtered call information. Upon receiving the target call archive file, the quality inspection server performs quality inspection analysis on it to obtain the analysis results for the filtered call information.

[0172] In one possible implementation, when the acquisition module 110 acquires the first raw audio stream, it specifically performs the following operations:

[0173] During the network call connection, the microphone component is used to collect the voice of the first party to obtain an analog voice electrical signal;

[0174] The analog voice signal is converted from analog to digital to obtain the first original audio stream.

[0175] In one possible implementation, the hybrid encoding module 130 is also used to perform the following operations:

[0176] Obtain the call information corresponding to the network call connection; the call information includes the account information of the first object account, the account information of the second object account, and the network call connection environment information;

[0177] The call information is processed by feature extraction to obtain the first object account features, the second object account features, and the network call connection environment features;

[0178] The current call scenario is obtained by performing call scenario prediction processing based on the characteristics of the first object account, the characteristics of the second object account, and the characteristics of the network call connection environment.

[0179] If the current call scenario is a recorded call scenario, then the step of performing audio mixing and encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection is executed.

[0180] The network call recording apparatus provided in this application embodiment allows the first terminal to obtain a first original audio stream when a network call connection is established between a first terminal and a second terminal. The first terminal then receives the transmitted audio stream sent by the second terminal through the network call connection, performs audio transmission decoding processing on the transmitted audio stream to obtain a second original audio stream, and finally performs audio mixing encoding processing on the first and second original audio streams to obtain a call recording file for the network call connection. The network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal. The first original audio stream refers to the audio stream generated by the first terminal based on the voice of a first object during the network call connection process. The first object refers to the object that logs into the first communication application in the first terminal through the first object's account. The second original audio stream refers to the audio stream generated by the second terminal based on the voice of a second object during the network call connection process. The second object refers to the object that logs into the second communication application in the second terminal through the second object's account. The method provided in this application embodiment uses a terminal to complete the recording of network calls, eliminating the need for dedicated call recording equipment, significantly reducing recording costs, and is no longer limited to a specific call system, greatly improving the versatility of network call recording.

[0181] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 8 As shown above, Figure 7 The network call recording device 1 in the corresponding embodiment can be applied to a computer device 1000, which may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 8 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0182] In such Figure 8 In the computer device 1000 shown, the network interface 1004 provides network communication elements; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0183] When the first terminal and the second terminal establish a network call connection through a communication application, the first original audio stream is obtained; the network call connection is established based on the first communication application in the first terminal and the second communication application in the second terminal; the first original audio stream refers to the audio stream generated by the first terminal based on the voice of the first object of the first object during the network call connection process; the first object refers to the object that logs in to the first communication application in the first terminal through the first object account;

[0184] The system receives the transmitted audio stream sent by the second terminal through a network call connection, performs audio transmission decoding processing on the transmitted audio stream, and obtains the second original audio stream. The second original audio stream refers to the audio stream generated by the second terminal based on the voice of the second object during the network call connection. The second object refers to the object that logs into the second communication application in the second terminal through the second object's account.

[0185] The first and second original audio streams are subjected to audio mixing and encoding to obtain a call recording file for the network call connection.

[0186] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 , Figure 5 The description of the data processing method in any corresponding embodiment will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0187] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned network call recording device 1. The computer program includes program instructions, and when the processor executes the program instructions, it can execute the aforementioned... Figure 3 , Figure 5The description of the data processing method in any corresponding embodiment is already provided, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.

[0188] The aforementioned computer-readable storage medium can be the network call recording device provided in any of the foregoing embodiments or the internal storage unit of the aforementioned computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0189] Furthermore, it should be noted that this application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned... Figure 3 , Figure 4 The method provided in any of the corresponding embodiments.

[0190] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0191] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0192] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the foregoing description as a network element. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described network elements using different methods for each specific application, but such implementation should not be considered beyond the scope of this application.

[0193] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A network call recording method, characterized by, The method is executed by a first terminal, and the method comprises: obtaining a first original audio stream when the first terminal establishes a network call connection with a second terminal; the network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal; the first original audio stream refers to an audio stream generated by the first terminal based on first object voice of a first object in the process of the network call connection; the first object refers to an object logging in the first communication application in the first terminal through a first object account; receiving a transmission audio stream sent by the second terminal through the network call connection, and performing audio transmission decoding processing on the transmission audio stream to obtain a second original audio stream; the second original audio stream refers to an audio stream generated by the second terminal based on second object voice of a second object in the process of the network call connection; the second object refers to an object logging in the second communication application in the second terminal through a second object account; performing audio mixing and encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

2. The method of claim 1, wherein, Further comprising: obtaining a recording notification broadcast file; the recording notification broadcast file is an audio file generated for recording notification broadcast voice; generating a mixed audio transmission audio stream according to the first original audio stream and the recording notification broadcast file; sending the mixed audio transmission audio stream to the second terminal, so that the second terminal plays the recording notification broadcast voice in the process of playing the first object voice according to the mixed audio transmission audio stream; when receiving a recording confirmation notification sent by the second terminal for the recording notification broadcast voice, performing the step of performing audio mixing and encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

3. The method of claim 2, wherein, The generating a mixed audio transmission audio stream according to the first original audio stream and the recording notification broadcast file comprises: generating an original recording notification broadcast file according to the recording notification broadcast file; generating a recording notification broadcast audio stream according to the recording notification broadcast file; the frame size of the audio frames contained in the recording notification broadcast audio stream is the same as the frame size of the audio frames contained in the first original audio stream; performing mixing processing on the first original audio stream and the recording notification broadcast audio stream to obtain a mixed audio stream; performing audio transmission encoding processing on the mixed audio stream to obtain a mixed audio transmission audio stream.

4. The method of claim 3, wherein, The generating an original recording notification broadcast file according to the recording notification broadcast file comprises: performing format identification processing on the recording notification broadcast file to obtain a format identification result; if the format identification result is a lossless compressed audio file format, the recording notification broadcast file is determined as an original recording notification broadcast file; if the format identification result is a lossy compressed audio file format, performing audio decoding processing on the recording notification broadcast file to obtain an original recording notification broadcast file.

5. The method of claim 3, wherein, The generating a recording notification broadcast audio stream according to the recording notification broadcast file comprises: Acquiring an audio frame size of a first original audio stream; According to the audio frame size, performing data reading processing on file data contained in the recording notification broadcast file to obtain N reading file data; the data amount of each reading file data is equal to the audio frame size; According to the audio frame format information of the first original audio stream, performing format conversion on the N reading file data respectively to obtain N recording notification broadcast audio frames; The N recording notification broadcast audio frames are determined as a recording notification broadcast audio stream.

6. The method of claim 5, wherein, The data reading processing on the file data contained in the recording notification broadcast file according to the audio frame size to obtain N reading file data includes: Generating a reading buffer according to the audio frame size; Starting to read the recording notification broadcast file from the position pointed to by the i-th moving file pointer to obtain i-th reading file data; the 1st moving file pointer points to the file start position of the recording notification broadcast file; If the data amount of the i-th reading file data is greater than the audio frame size, sequentially writing the i-th reading file data into the reading buffer, determining the data contained in the reading buffer as the i-th reading file data, and emptying the reading buffer; According to the audio frame size, moving the i-th moving file pointer backward in the recording notification broadcast file to obtain an i+1-th backward moving file pointer, and continuing to read the recording notification broadcast file starting from the position pointed to by the i+1-th backward moving file pointer; If the data amount of the i-th reading file data is less than or equal to the audio frame size, sequentially writing the i-th reading file data into the reading buffer, performing default data padding processing on the remaining buffer area of the reading buffer, determining the data contained in the reading buffer as the i-th reading file data, and determining the first i reading file data as the N reading file data.

7. The method of claim 3, wherein, The mixing processing on the first original audio stream and the recording notification broadcast audio stream to obtain a mixed audio stream includes: Performing addition processing on the first original audio stream and the recording notification broadcast audio stream to obtain a sum audio stream; Performing normalization processing on the sum audio stream to obtain a mixed audio stream.

8. The method of claim 1, wherein, The audio mixing encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection includes: Performing left channel conversion processing on the first original audio stream to obtain a left channel audio stream; Performing right channel conversion processing on the second original audio stream to obtain a right channel audio stream; Performing audio stream merging processing on the left channel audio stream and the right channel audio stream to obtain a dual-channel audio stream; Performing audio encoding processing on the dual-channel audio stream to obtain a call recording file for the network call connection.

9. The method of claim 8, wherein, The audio stream merging processing on the left channel audio stream and the right channel audio stream to obtain a dual-channel audio stream includes: acquiring a jth left channel audio frame in the left channel audio stream and a jth right channel audio frame in the right channel audio stream in a traversal manner; the jth left channel audio frame and the jth right channel audio frame have a time alignment relationship; j is a positive integer; performing an alternating merging process on M sample points contained in the target left channel audio frame and M sample points contained in the target right channel audio frame, to obtain a jth dual-channel audio frame containing 2M sample points; M is a positive integer; generating a dual-channel audio stream according to the obtained dual-channel audio frames when the traversal of the left channel audio frames contained in the left channel audio stream is completed.

10. The method of claim 1, wherein, Further comprising: when detecting that the network call connection is disconnected, acquiring call information corresponding to the network call connection; generating a call archive file according to the call information and the call recording file, and sending the call archive file to an archive server, so that the archive server stores the call archive file; the archive server is configured to send a target call archive file to a quality inspection server when receiving a file acquisition request containing screening call information sent by the quality inspection server; the target call archive file refers to a call archive file containing call information matching the screening call information; the quality inspection server is configured to perform quality inspection analysis processing on the target call archive file when receiving the target call archive file, to obtain a quality inspection analysis result for the screening call information.

11. The method of claim 1, wherein, acquiring a first original audio stream, comprising: acquiring a first object voice through a microphone assembly during the network call connection process to obtain an analog voice electrical signal; performing analog-digital conversion processing on the analog voice electrical signal to obtain a first original audio stream.

12. The method of claim 1, wherein, Further comprising: acquiring call information corresponding to the network call connection; the call information contains account information of the first object account, account information of the second object account, and network call connection environment information; performing feature extraction processing on the call information to obtain first object account features, second object account features, and network call connection environment features; performing call scene prediction processing according to the first object account features, the second object account features, and the network call connection environment features to obtain a current call scene; if the current call scene belongs to a recording call scene, performing the audio mixing and encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

13. A network conversation recording apparatus characterized by comprising: The device is run by a first terminal, and the device comprises: an acquisition module configured to acquire a first original audio stream when the first terminal establishes a network call connection with a second terminal; the network call connection is established based on a first communication application in the first terminal and a second communication application in the second terminal; the first original audio stream refers to an audio stream generated by the first terminal based on a first object voice of a first object during the network call connection process; the first object refers to an object logging into the first communication application in the first terminal through a first object account; The decoding module is configured to receive, by the first terminal, a transmission audio stream sent by the second terminal through the network call connection, and perform audio transmission decoding processing on the transmission audio stream to obtain a second original audio stream; the second original audio stream refers to an audio stream generated by the second terminal based on second object voice of a second object in the network call connection process; the second object refers to an object that logs in a second communication application in the second terminal through a second object account; The mixed encoding module is configured to perform audio mixed encoding processing on the first original audio stream and the second original audio stream to obtain a call recording file for the network call connection.

14. A computer device, comprising: Comprise: A processor, a memory and a network interface; The network interface is configured to provide a data communication function, the memory is configured to store program code, and the processor is configured to call the program code to execute the method in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer program stored in the computer readable storage medium is adapted to be loaded and executed by the processor to execute the method in any one of claims 1-12.

16. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to execute the method in any one of claims 1-12.