Audio stream processing method based on RTSP protocol and related device
By verifying the RTSP protocol support of the voiceprint device, the audio streaming is optimized by using summary authentication and RTSP SETUP commands, combining multi-threaded processing and dynamic port allocation, the compatibility, security and efficiency problems in power inspection audio streaming are solved, and the reliability and adaptability of the system are improved.
Patent Information
- Application Number
- CN202510711793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-19
AI Technical Summary
The existing technology has problems such as lack of dynamic adaptation capabilities, low data analysis efficiency, insufficient multi-thread processing capabilities, weak authentication security mechanism, inefficient port resource management and lack of dynamic distribution capabilities in the audio streaming architecture in the field of power inspection, which affects the reliability, security and scalability of the voiceprint monitoring system.
Verify the protocol support of the voiceprint device by sending RTSP GET commands, enhance transmission security by using digest authentication, sending RTSP SETUP commands to specify transmission parameters, receive and parse RTP audio packets, push preset format audio streams to the client, and combine multi-threaded processing and dynamic port allocation to achieve efficient transmission and dynamic distribution of audio streams.
It improves the compatibility and deployment efficiency of multi-vendor equipment, ensures the security and real-time nature of data transmission, meets the differentiated monitoring needs of power inspection scenarios, and improves the flexibility and adaptability of the system.
Smart Images

Figure CN120512423A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of streaming media transmission and audio processing, relates to power inspection audio stream processing technology, and in particular to an audio stream processing method based on RTSP protocol and related devices. Background Art
[0002] In the field of power inspection, the equipment status monitoring system based on voiceprint perception and intelligent analysis technology is becoming a key technology to ensure the full operation of the power grid; by deploying high-precision voiceprint monitoring terminals on key equipment in substations, real-time, contactless collection of operating audio streams of key equipment such as transformers, circuit breakers and capacitors can be achieved, providing multi-dimensional data support for equipment status assessment and fault warning.
[0003] At present, the audio stream in the field of power inspection mainly relies on RTSP (Real-Time Streaming Protocol) to build a transmission control system. Its core process covers standardized command interaction and underlying data transmission mechanism. Specifically, the client initiates a sequence of commands such as OPTIONS, DESCRIBE, SETUP or PLAY based on preset device capability parameters (such as encoding format, resolution and port number) to complete session negotiation and dynamic establishment of media stream transmission channels. In RTP (Real-time Transport Protocol), the client sends a series of commands such as OPTIONS, DESCRIBE, SETUP or PLAY to complete session negotiation and dynamic establishment of media stream transmission channels. The real-time transport protocol (RTP) data layer is used to reassemble packets through serial number association and timestamp calibration mechanisms to ensure the timing alignment and integrity verification of audio stream data. The original PCM audio data is converted into formats (such as MP3 / AAC encoding) using tools such as FFmpeg to meet the compatibility requirements of terminal players. The above solution has been applied on a large scale in the field of power inspection. The substation voiceprint equipment collects the operating audio streams of key equipment such as transformers and circuit breakers in real time, and combines RTP protocol parsing and multi-format transcoding technology to provide standardized data input for the voiceprint feature analysis algorithm, thereby realizing online diagnosis and early warning of potential faults such as mechanical looseness and partial discharge.
[0004] However, the current audio stream transmission architecture still has defects such as lack of dynamic adaptation capabilities, low data parsing efficiency, insufficient multi-threaded processing capabilities, weak authentication security mechanisms, inefficient port resource management and lack of dynamic distribution capabilities, which directly restrict the reliability, security and scalability of voiceprint monitoring systems in complex industrial environments.
[0005] The specific analysis is as follows: Lack of dynamic adaptation capabilities: Traditional RTSP clients rely on manual pre-configuration of device parameters and are unable to automatically detect the device's supported protocol version and encoding capabilities, resulting in poor compatibility with multi-vendor devices and inefficient deployment. Inefficient data parsing: RTP packet reassembly logic uses static rule matching and lacks a dynamic optimization mechanism based on network status. This can easily lead to packet disarray and loss in complex network environments, significantly reducing real-time performance. Inadequate multi-threaded processing capabilities: The single-threaded architecture performs audio stream processing and feature extraction tasks serially, failing to meet the millisecond-level response latency requirements of power inspection scenarios and hindering the timeliness of fault diagnosis. Weak authentication security mechanisms: Some devices still use Basic authentication with plaintext transmission, posing the risk of man-in-the-middle attacks and failing to meet the data confidentiality and integrity security requirements of the Industrial Internet. Inefficient port resource management: Fixed UDP port configurations are prone to port conflicts and insufficient port utilization in scenarios with multiple devices running concurrently, resulting in wasted resources and reduced transmission efficiency. Lack of dynamic distribution capabilities: The traditional architecture adopts a "pull-type" data transmission mode, which cannot push specific audio streams in real time based on user subscription requirements or changes in device status, and is difficult to adapt to the demand for differentiated monitoring data in power inspections. Summary of the Invention
[0006] In response to the technical problems existing in the prior art, the present invention provides an audio stream processing method and related devices based on the RTSP protocol to solve the technical problems that still exist in the current audio stream transmission architecture, such as lack of dynamic adaptation capability, low data parsing efficiency, insufficient multi-threaded processing capability, weak authentication security mechanism, inefficient port resource management and lack of dynamic distribution capability.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is: The present invention provides an audio stream processing method based on the RTSP protocol, which is used for a streaming media server. The audio stream processing method based on the RTSP protocol includes: Receive the audio stream playback request sent by the client; Send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming; Send a digest authentication request to the voiceprint device and receive the session description protocol information returned by the voiceprint device; Send the RTSP SETUP command to the voiceprint device; the RTSP SETUP command is used to specify the transmission parameters of the audio stream; Receive, parse and reassemble the RTP audio data packets sent by the voiceprint device to obtain the audio stream in the preset format; Pushes audio streams in a preset format to the client.
[0008] Furthermore, the audio stream playback request includes the voiceprint device ID, the user IP address and the audio stream format requirement.
[0009] Furthermore, the digest authentication request includes a user name and password encrypted using the MD5 encryption algorithm; the session description protocol information includes an encoding format and a sampling rate of the audio stream.
[0010] Furthermore, after sending the digest authentication request to the voiceprint device and receiving the session description protocol information returned by the voiceprint device, the following steps are also included: Query the UDP port list of the streaming media server and select an idle UDP port; wherein, the idle UDP port is used to receive the RTP audio data packet sent by the voiceprint device.
[0011] Furthermore, the audio stream in the preset format is an audio stream in the PCM format.
[0012] Furthermore, the process of pushing the audio stream in a preset format to the client includes: According to the audio stream playback request sent by the client, the audio stream in a preset format is pushed to the client via the UDP protocol or the TCP protocol; wherein, after receiving the audio stream in the preset format, the client plays it in real time.
[0013] The present invention also provides an audio stream processing system based on the RTSP protocol, which is used for a streaming media server. The audio stream processing system based on the RTSP protocol includes: The playback request receiving module is used to receive the audio stream playback request sent by the client; A protocol verification module is used to send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming; An authentication module, configured to send a digest authentication request to a voiceprint device and receive session description protocol information returned by the voiceprint device; The transmission channel establishment module is used to send the RTSP SETUP command to the voiceprint device; wherein the RTSP SETUP command is used to specify the transmission parameters of the audio stream; The audio receiving and processing module is used to receive, parse and reassemble the RTP audio data packets sent by the voiceprint device to obtain the audio stream in a preset format; The audio push module is used to push the audio stream in a preset format to the client.
[0014] The present invention also provides an electronic device, comprising: a processor suitable for executing a computer program; A computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the audio stream processing method based on the RTSP protocol is executed.
[0015] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the audio stream processing method based on the RTSP protocol is implemented.
[0016] The present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the audio stream processing method based on the RTSP protocol is implemented.
[0017] Compared with the prior art, the present invention has the following beneficial effects: The audio stream processing method based on RTSP protocol provided by the present invention verifies the protocol support of the voiceprint device by sending the RTSP GET command, uses digest authentication to enhance transmission security, specifies transmission parameters with the RTSP SETUP command to ensure data transmission accuracy, receives and parses the RTP data packet to obtain the preset format audio stream and pushes it to the client, effectively improving the compatibility and deployment efficiency of multi-vendor equipment, ensuring data transmission security, and laying the foundation for subsequent optimization of data analysis, realization of multi-threaded processing and dynamic distribution, etc., and is suitable for the high-reliability audio stream processing needs of power inspection scenarios; specifically, by sending RTSP The GET command is sent to the voiceprint device to verify whether it supports the RTSP protocol for streaming and to specify the audio stream transmission parameters. It can automatically adapt to the capabilities of different devices without the need for manual pre-configuration one by one, effectively improving the compatibility of multi-vendor devices and significantly improving deployment efficiency; receiving, parsing and reassembling the RTP audio data sent by the voiceprint device creates conditions for introducing a dynamic optimization mechanism based on network status; it can dynamically adjust the packet subassembly and reassembly rules according to the real-time status of the network to avoid problems such as packet disorder and packet loss in complex network environments, thereby improving the real-time and accuracy of data analysis; by sending a summary authentication request to the voiceprint device, it can effectively prevent man-in-the-middle attacks, meet the security requirements of the industrial Internet for data confidentiality and integrity, and ensure the security of audio stream data during transmission; by receiving client requests and pushing audio streams to the client, it can push specific audio streams in real time according to user subscription requirements or device status changes, meet the needs of power inspections for differentiated monitoring data, and improve the flexibility and adaptability of the system.
[0018] Furthermore, by querying the UDP port list of the streaming media server and selecting an idle UDP port to receive the RTP audio data packet sent by the voiceprint device, resource optimization is achieved through the dynamic port allocation strategy.
[0019] The RTSP-based audio stream processing system, electronic device, computer-readable storage medium, and computer program product provided by the present invention have all the advantages of the RTSP-based audio stream processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 Flowchart of the audio stream processing method based on the RTSP protocol provided in Example 1; Figure 2 A block diagram of the audio stream processing system based on the RTSP protocol provided in Example 2; Figure 3 This is a structural block diagram of the electronic device provided in Example 3. DETAILED DESCRIPTION
[0022] In order to make the technical problems, technical solutions, and beneficial effects solved by this application more clearly understood, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application; it is obvious that the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of this application.
[0023] The present invention provides an audio stream processing method based on the RTSP protocol, which is used for a streaming media server. The audio stream processing method based on the RTSP protocol comprises the following steps: Step 100: Receive an audio stream playback request sent by a client.
[0024] Step 200: Send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming.
[0025] Step 300: Send a digest authentication request to the voiceprint device, and receive session description protocol information returned by the voiceprint device.
[0026] Step 400: Send an RTSP SETUP command to the voiceprint device; wherein the RTSP SETUP command is used to specify transmission parameters of the audio stream.
[0027] Step 500: Receive, parse and reassemble the RTP audio data packet sent by the voiceprint device to obtain an audio stream in a preset format.
[0028] Step 600: Push the audio stream in a preset format to the client.
[0029] The audio stream processing method based on the RTSP protocol described in the present invention verifies the protocol support of the voiceprint device by sending the RTSP GET command, thereby enhancing the dynamic adaptation capability, helping to be compatible with multi-vendor equipment and improving deployment efficiency; sending a summary authentication request ensures the security of data transmission and meets the security requirements of the industrial Internet; sending the RTSP SETUP command to specify the transmission parameters, combined with the subsequent reception and parsing of the RTP data packet to obtain a preset format audio stream and push it to the client, optimizes the data parsing and transmission process, and can improve data parsing efficiency. Its overall architecture is more conducive to multi-threaded processing to meet low latency requirements, while solving the problem of inefficient port resource management. It also realizes dynamic distribution through push mode, adapts to the differentiated monitoring data requirements of power inspections, and thereby improves the reliability, security and scalability of the voiceprint monitoring system in complex industrial environments.
[0030] The following further explains the audio stream processing method based on the RTSP protocol provided by the present invention with some specific embodiments: Example 1 As attached Figure 1 As shown, this embodiment 1 provides an audio stream processing method based on the RTSP protocol, including the following steps: Step 1: Receive an audio stream playback request from the client. Specifically, the user sends a request to play an audio stream to the streaming media server through the client's preset application. The client, for example, is a dedicated mobile terminal for power inspection or monitoring center software. The audio stream playback request includes the identification information of the target voiceprint device, such as the voiceprint device ID, the user's IP address, and the audio stream format requirements. The audio stream format requirements include the audio sampling rate and encoding format to ensure that the streaming media server provides the user with an audio stream that meets the user's needs.
[0031] When the streaming media server receives the audio stream playback request sent by the client, it will record the subscription information of the audio stream playback request in the internal database or memory structure of the streaming media server to ensure that the audio stream can be accurately pushed to the corresponding user in the future, and the audio stream can be processed and converted accordingly according to the audio format requirements specified by the user.
[0032] Step 2: Create an independent thread. The independent thread created is used to process the audio stream, while the main thread of the streaming media server continues to respond to other requests. Specifically, after receiving a play request, the streaming media server will immediately create a new independent thread to handle the user's audio stream reception and distribution tasks. This thread is responsible for obtaining audio data from the voiceprint device, converting it into a playable format, and then pushing it to the user.
[0033] It should be noted that after creating an independent thread, the main thread of the streaming media server continues to run to respond to various requests that may be initiated by other users, such as new playback requests and device control requests; the design of a multi-threaded architecture can fully utilize the computing resources of the streaming media server, improve concurrent processing capabilities, and ensure that multiple users can play audio streams at the same time without interfering with each other.
[0034] Step 3: Send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming.
[0035] Specifically, the steps are as follows: Use the independent thread of the streaming media server to send the RTSP GET command to the voiceprint device to query whether the voiceprint device supports RTSP protocol streaming; after receiving the RTSP GET command, the voiceprint device returns a response status code to the streaming media server; if the response status code is 200, it means that the voiceprint device supports RTSP real-time streaming function, and the streaming media server can continue the subsequent operation process; if the response status code is not 200, it means that the voiceprint device does not support RTSP real-time streaming function, and the streaming media server will prompt the user that the device is incompatible or try to use other protocols to connect.
[0036] Step 4: Send a digest authentication request to the voiceprint device and receive the session description protocol information returned by the voiceprint device. The digest authentication request includes the username and password encrypted with the MD5 algorithm; the session description protocol information includes the encoding format and sampling rate of the audio stream.
[0037] Specifically, the process is as follows: After confirming that the voiceprint device supports RTSP protocol streaming, the streaming media server sends an RTSP DESCRIBE command to the voiceprint device. The RTSP DESCRIBE command is used to obtain the SDP (Session Description Protocol) information of the voiceprint device. The SDP information of the voiceprint device includes the encoding format, sampling rate and transmission protocol of the audio stream, which is used to subsequently establish the audio stream transmission channel.
[0038] In order to ensure that only authorized users can access the audio stream of the voiceprint device, when the streaming media server sends the RTSP DESCRIBE command to the voiceprint device, the username and password encrypted by the MD5 encryption algorithm are added to the RTSP DESCRIBE command to form a complete digest authentication request; when the voiceprint device receives the digest authentication request, it uses the same encryption algorithm to decrypt and verify the received username and password; if the verification is successful, the voiceprint device will return the SDP information; if the verification fails, the voiceprint device will reject the streaming media server's request and return a corresponding error message.
[0039] Step 5: Query the UDP port list of the streaming media server and select an idle UDP port; wherein, the idle UDP port is used to receive the RTP audio data packet sent by the voiceprint device.
[0040] It should be noted that the streaming media server queries the list of available UDP ports on the current server to determine which ports are idle for receiving RTP audio data packets from the voiceprint device; the process of querying the list of available UDP ports on the current server is implemented by querying the server's network configuration information or using a preset port management tool; after obtaining the available port list, an idle UDP port is selected according to the preset port selection strategy; for example, an idle port in the range of 5004-5100 is selected for subsequent establishment of an RTP transmission channel with the voiceprint device to receive audio data packets.
[0041] Step 6: Send an RTSP SETUP command to the voiceprint device. The RTSP SETUP command is used to specify the transmission parameters of the audio stream. In the RTSP SETUP command, the streaming server pre-specifies the transmission parameters, including the transmission protocol used and the idle UDP port, to establish the transmission session channel for the audio stream.
[0042] After receiving the RTSP SETUP command, the voiceprint device begins to verify and process the transmission parameters in the RTSP SETUP command; if the transmission parameters are legal and the voiceprint device can support them, the voiceprint device will return a confirmation message to the streaming media server, indicating that it agrees to establish a session channel; at this point, the audio stream transmission session channel between the voiceprint server and the voiceprint device is officially established, preparing for subsequent audio data transmission.
[0043] Step 7: Receive, parse, and reassemble the RTP audio data packets sent by the voiceprint device to obtain an audio stream in a preset format, wherein the audio stream in the preset format is an audio stream in PCM format.
[0044] Specifically, the steps are as follows: Step 71: The streaming media server starts to receive RTP audio data packets from the voiceprint device through the previously established audio stream transmission session channel.
[0045] Step 72: After receiving the RTP data packet, the streaming media server parses the header information of the RTP data packet; the header information of the RTP data packet includes a timestamp and a sequence number; the timestamp is used to identify the time sequence of the audio data to ensure that the audio can be played in the correct time sequence; the sequence number is used to detect the loss and disorder of the RTP data packet so that the streaming media server can perform corresponding processing.
[0046] Step 73: After the header information is parsed, the audio payload data is extracted from the RTP data packet; wherein the audio payload data is the audio encoding data actually carried in the RTP data packet, which includes the compressed or encoded audio information.
[0047] Step 74: Decode and reassemble the extracted audio payload data to convert it into continuous PCM (Pulse Code Modulation) format audio frames to obtain a PCM format audio stream; wherein the PCM format is an uncompressed audio format that can be directly recognized and played by an audio playback device.
[0048] Step 8: Push the audio stream in the preset format to the client. Specifically, according to the audio stream playback request sent by the client, the audio stream in the preset format is pushed to the client via the UDP protocol or the TCP protocol; wherein, after receiving the audio stream in the preset format, the client plays it in real time.
[0049] It should be noted that after the client receives the PCM format audio stream pushed by the streaming server, it starts to call the local pcmplayer player or other audio player that supports PCM format playback to prepare to play the audio; the pcmplayer player decodes the received PCM format audio stream, and then plays the decoded audio data in real time, allowing the user to hear the audio information from the voiceprint device.
[0050] It should be noted that when the voiceprint device does not support RTSP protocol streaming, the RTSP protocol can be replaced with HTTP Live Streaming (HLS) protocol. At this time, steps 3-6 are synchronously adjusted to segmented requests; specifically, the GET, DESCRIBE, and SETUP commands of the RTSP protocol are replaced with segmented requests initiated through the HTTP protocol; at this time, the streaming media server requests the M3U8 index file of the audio file from the voiceprint device, and then requests each TS file segment in turn according to the information in the index file, thereby realizing the transmission of the audio stream.
[0051] It should also be noted that the username and password encrypted by the MD5 encryption algorithm in step 4 can be replaced with the username and password encrypted and authenticated by HMAC-SHA256 to further improve the security of the system.
[0052] The audio stream processing method based on the RTSP protocol described in this embodiment 1 has obvious gain effects on improving the efficiency of fault tracing of power inspection services, optimizing resource utilization, and strengthening security protection. Specifically, by multi-threaded parallel processing of audio stream reception, decoding and distribution tasks, the end-to-end transmission delay is significantly reduced. Among them, the delay in a typical scenario is less than 200ms, which meets the millisecond-level abnormal alarm requirements of the voiceprint of power equipment. For example, the transient audio characteristics of loose transformer core or partial discharge can be captured in real time. The MD5 dynamic digest authentication mechanism is used to log in to the voiceprint collection device, and combined with the power intranet whitelist strategy, the risk of plain text password transmission is prevented, and the security of equipment access in unmanned outdoor scenarios such as substations and transmission towers is guaranteed. Based on the dynamic allocation of UDP port resources, it supports the concurrent access of hundreds of voiceprint sensors in the substation, avoids the address conflicts caused by traditional fixed port configuration, and adapts to the heterogeneous network environment of the power Internet of Things.
[0053] In this embodiment 1, targeted distribution of audio streams can be achieved through the subscription-publish mode. Specified audio streams can be pushed to mobile inspection terminals, regional monitoring centers, or AI analysis servers based on the power inspection business logic (such as equipment type and regional affiliation), supporting multi-department collaborative fault diagnosis. A parsing engine for commonly used formats for power voiceprint monitoring, such as PCM / G.711, is built in, supporting direct storage and playback of original audio streams on the edge side, eliminating the computational overhead of the transcoding link, and adapting to resource-constrained scenarios at the edge nodes of power inspections.
[0054] Example 2 As attached Figure 2 As shown, this embodiment 2 provides an audio stream processing system based on the RTSP protocol, which is used for a streaming media server. The audio stream processing system based on the RTSP protocol includes a playback request receiving module, a thread module, a protocol verification module, an authentication module, a port query module, a transmission channel establishment module, an audio receiving processing module and an audio push module.
[0055] A play request receiving module for receiving an audio stream play request sent by a client; a thread module for creating an independent thread; a protocol verification module for sending an RTSP GET command to a voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol stream acquisition; an authentication module for sending a summary authentication request to the voiceprint device and receiving session description protocol information returned by the voiceprint device; a port query module for querying the UDP port list of the streaming media server and selecting an idle UDP port; wherein the idle UDP port is used to receive the RTP audio data packet sent by the voiceprint device; a transmission channel establishment module for sending an RTSP SETUP command to the voiceprint device; wherein the RTSP SETUP command is used to specify the transmission parameters of the audio stream; an audio receiving and processing module for receiving, parsing and reassembling the RTP audio data packet sent by the voiceprint device to obtain an audio stream in a preset format; an audio push module for pushing an audio stream in a preset format to the client.
[0056] Example 3 As attached Figure 3 As shown, this embodiment 3 provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of the audio stream processing method based on the RTSP protocol when executing the computer program. Alternatively, the processor implements the functions of each module in the above-mentioned audio stream processing system based on the RTSP protocol when executing the computer program.
[0057] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing preset functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0058] The electronic device may be a computing device such as a desktop computer, laptop, PDA, or cloud server. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the above are examples of electronic devices and do not constitute a limitation on electronic devices. The electronic device may include more components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0059] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.
[0060] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0061] The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback); the data storage area may store data generated based on the use of the mobile phone (such as audio data and a phone book). Furthermore, the memory may include high-speed random access memory (RAM) and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0062] Example 4 This embodiment 4 further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the audio stream processing method based on the RTSP protocol are implemented.
[0063] If the modules / units integrated in the RTSP protocol-based audio stream processing system are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0064] Based on this understanding, the present invention implements all or part of the processes in the above-mentioned audio stream processing method based on the RTSP protocol, and can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the above-mentioned audio stream processing method based on the RTSP protocol. The computer program includes computer program code, which can be in source code form, object code form, executable file, or preset intermediate form.
[0065] The computer-readable storage medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0066] Example 5 This embodiment 5 provides a computer product, which includes a computer program product, which is stored in a computer-readable storage medium; the processor of the electronic device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the electronic device can execute the audio stream processing method based on the RTSP protocol described in embodiment 1, which will not be repeated here.
[0067] It should be noted that a person skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0068] The above embodiment is only one of the implementation methods that can realize the technical solution of the present invention. The scope of protection claimed by the present invention is not limited only to this embodiment, but also includes changes, replacements and other implementation methods that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention.
Claims
1. A method for processing audio streams based on RTSP protocol, characterized in that: For a streaming media server; the audio stream processing method based on the RTSP protocol includes: Receive the audio stream playback request sent by the client; Send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming; Send a digest authentication request to the voiceprint device and receive the session description protocol information returned by the voiceprint device; Send the RTSP SETUP command to the voiceprint device; the RTSP SETUP command is used to specify the transmission parameters of the audio stream; Receive, parse and reassemble the RTP audio data packets sent by the voiceprint device to obtain the audio stream in the preset format; Pushes audio streams in a preset format to the client.
2. The method for processing an audio stream based on RTSP protocol according to claim 1, wherein: The audio stream playback request includes the voiceprint device ID, user IP address, and audio stream format requirements.
3. The method for processing an audio stream based on RTSP protocol according to claim 1, wherein: The digest authentication request includes the user name and password encrypted using the MD5 encryption algorithm; the session description protocol information includes the encoding format and sampling rate of the audio stream.
4. The method for processing an audio stream based on RTSP protocol according to claim 1, wherein: After sending a digest authentication request to the voiceprint device and receiving the session description protocol information returned by the voiceprint device, it also includes: Query the UDP port list of the streaming media server and select an idle UDP port; wherein, the idle UDP port is used to receive the RTP audio data packet sent by the voiceprint device.
5. The method for processing an audio stream based on RTSP protocol according to claim 1, wherein: The default audio stream format is PCM format audio stream.
6. The method for processing an audio stream based on RTSP protocol according to claim 1, wherein: The process of pushing a preset audio stream to the client includes: According to the audio stream playback request sent by the client, the audio stream in a preset format is pushed to the client via the UDP protocol or the TCP protocol; wherein, after receiving the audio stream in the preset format, the client plays it in real time.
7. An audio stream processing system based on RTSP protocol, characterized in that: For a streaming media server, the RTSP-based audio stream processing system includes: The playback request receiving module is used to receive the audio stream playback request sent by the client; A protocol verification module is used to send an RTSP GET command to the voiceprint device; wherein the RTSP GET command is used to verify whether the voiceprint device supports RTSP protocol streaming; An authentication module, configured to send a digest authentication request to a voiceprint device and receive session description protocol information returned by the voiceprint device; The transmission channel establishment module is used to send the RTSP SETUP command to the voiceprint device; wherein the RTSP SETUP command is used to specify the transmission parameters of the audio stream; The audio receiving and processing module is used to receive, parse and reassemble the RTP audio data packets sent by the voiceprint device to obtain the audio stream in a preset format; The audio push module is used to push the audio stream in a preset format to the client.
8. An electronic device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the audio stream processing method based on the RTSP protocol according to any one of claims 1 to 6 is executed.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the audio stream processing method based on the RTSP protocol is implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the audio stream processing method based on the RTSP protocol is implemented according to any one of claims 1 to 6.