Voice intercom method and device of camera, computer device and storage medium
By defining the intercom mode of the cameras in the video surveillance platform and adopting a standardized communication protocol, the compatibility issues of cameras from different manufacturers are resolved, the configuration process is simplified, and the stability and compatibility of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2024-11-19
- Publication Date
- 2026-04-28
AI Technical Summary
The inconsistent implementation rules of the intercom function of cameras from different manufacturers lead to compatibility issues when integrating third-party cameras into video surveillance platforms, increasing system complexity and maintenance costs.
The target platform's server sends a data packet carrying a target identifier to the camera. If a return data packet is received, the first intercom mode is determined; otherwise, a target address information data packet is sent to determine the second intercom mode. Audio data is then sent based on the matching mode, using either RTP or TCP protocols for communication.
It simplifies the camera configuration process on the target platform, improves platform compatibility, reduces manual intervention and configuration errors, enhances system reliability and stability, and reduces operation and maintenance costs.
Smart Images

Figure CN119496867B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a voice intercom method, apparatus, computer device, computer-readable storage medium, and computer program product for a camera. Background Technology
[0002] With the development of internet technology, video surveillance technology has emerged and is widely used in fields such as security, management, and transportation to monitor and record on-site situations in real time. As technology advances, video surveillance systems not only need to provide high-quality video images but also need to implement two-way voice communication capabilities to enable real-time voice communication and command in emergency situations.
[0003] However, different manufacturers have different implementation rules for the intercom function of cameras, which leads to compatibility issues in practical applications. As a result, when video surveillance platforms incorporate a large number of third-party surveillance cameras, they need to perform complex adaptation and configuration work, which increases the complexity and maintenance cost of the system. Summary of the Invention
[0004] Therefore, it is necessary to provide a camera voice intercom method, device, computer equipment, computer-readable storage medium, and computer program product that can simplify the camera configuration process and improve the compatibility of platforms connecting to the camera, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a voice intercom method for a camera, comprising:
[0006] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0007] If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port;
[0008] If the return data packet carrying the target identifier is not received at the first port, a second data packet carrying the target address information is sent to the camera. If the return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0009] Based on the intercom mode matched by the camera, audio data packaged with the target communication protocol is sent to the camera so that the camera receives the audio data and plays the voice corresponding to the audio data.
[0010] In one embodiment, when a camera is connected to a target platform, the step of the server corresponding to the target platform sending a first data packet carrying a target identifier to the camera includes:
[0011] When the camera is connected to the target platform, the server generates the target identifier;
[0012] Create a session establishment request packet carrying the target identifier;
[0013] The session establishment request data packet is sent to the camera via the SIP protocol.
[0014] In one embodiment, the target identifier is the synchronization source identifier.
[0015] In one embodiment, the step of sending a second data packet carrying target address information to the camera includes:
[0016] Create a session establishment request data packet carrying the target address information;
[0017] The session establishment request data packet is sent to the camera via the SIP protocol;
[0018] The step of receiving the return data packet sent by the camera based on the target address information includes:
[0019] The camera received a TCP connection establishment request data packet carrying the target address information.
[0020] In one embodiment, the steps following the receipt of the TCP connection establishment request data packet carrying the target address information from the camera include:
[0021] The address of the camera is obtained based on the TCP connection establishment request data packet;
[0022] A separate target intercom port is assigned to the camera to send the audio data to the address of the camera via the target intercom port.
[0023] In one embodiment, the audio payload data corresponding to the audio data consists of an integer number of audio encoded frames, and the audio bitstream corresponding to the audio data conforms to the target encoding standard.
[0024] In one embodiment, the step of sending audio data packaged with a target communication protocol to the camera, so that the camera receives the audio data and plays the corresponding speech, includes:
[0025] Obtain the device type of the camera;
[0026] The data format of the audio data is determined based on the device type;
[0027] The system sends audio data in the specified data format, packaged via the RTP protocol, to the camera so that the camera can play the corresponding audio data when it receives the audio data.
[0028] Secondly, this application also provides a camera-based voice intercom device, comprising:
[0029] The first determining module is used to send a first data packet carrying a target identifier to the camera when the camera is connected to the target platform; if a return data packet carrying the target identifier is received through the first port, the intercom mode matched by the camera is determined to be the first intercom mode, and the first intercom mode communicates with the camera through the first port.
[0030] The second determining module is configured to send a second data packet carrying target address information to the camera if the return data packet carrying the target identifier is not received at the first port, and if the return data packet sent by the camera based on the target address information is received, determine that the intercom mode matched by the camera is the second intercom mode, and the second intercom mode communicates with the camera through the second port.
[0031] The intercom module is used to send audio data packaged with a target communication protocol to the camera based on the intercom mode matched by the camera, so that the camera receives the audio data and plays the voice corresponding to the audio data.
[0032] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0033] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0034] If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port;
[0035] If the return data packet carrying the target identifier is not received at the first port, a second data packet carrying the target address information is sent to the camera. If the return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0036] Based on the intercom mode matched by the camera, audio data packaged with the target communication protocol is sent to the camera so that the camera receives the audio data and plays the voice corresponding to the audio data.
[0037] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0038] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0039] If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port;
[0040] If the return data packet carrying the target identifier is not received at the first port, a second data packet carrying the target address information is sent to the camera. If the return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0041] Based on the intercom mode matched by the camera, audio data packaged with the target communication protocol is sent to the camera so that the camera receives the audio data and plays the voice corresponding to the audio data.
[0042] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0043] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0044] If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port;
[0045] If the return data packet carrying the target identifier is not received at the first port, a second data packet carrying the target address information is sent to the camera. If the return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0046] Based on the intercom mode matched by the camera, audio data packaged with the target communication protocol is sent to the camera so that the camera receives the audio data and plays the voice corresponding to the audio data.
[0047] The aforementioned camera-based voice intercom method, device, computer equipment, computer-readable storage medium, and computer program product, when the camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera. If a return data packet carrying a target identifier is received through a first port, the intercom mode matched by the camera is determined to be the first intercom mode. If no return data packet carrying a target identifier is received at the first port, a second data packet carrying target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. In this way, after the camera is connected to the target platform, it attempts to determine which of the two intercom modes the camera supports. Compared to obtaining detailed information about the camera for complex adaptation and configuration of each camera, this simplifies the configuration process required for the camera to connect to the target platform and enables adaptation of intercom modes for cameras with different sound fields from different manufacturers. This improves platform compatibility, reduces the possibility of manual intervention and configuration errors, improves the reliability and stability of the target platform, and reduces operation and maintenance costs. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a diagram illustrating the application environment of a camera-based voice intercom method in one embodiment.
[0050] Figure 2 This is a flowchart illustrating a voice intercom method using a camera in one embodiment;
[0051] Figure 3This is a flowchart illustrating the first intercom mode matching step in one embodiment;
[0052] Figure 4 This is a flowchart illustrating the second intercom mode matching step in one embodiment;
[0053] Figure 5 This is a flowchart illustrating the voice intercom steps in the first intercom mode of one embodiment.
[0054] Figure 6 This is a flowchart illustrating the voice intercom steps in the second intercom mode of one embodiment.
[0055] Figure 7 This is a structural block diagram of the voice intercom device of a camera in one embodiment;
[0056] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0058] The camera voice intercom method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. The terminal and server constitute the target platform. The target platform can connect to multiple cameras 106 for unified management. When a camera 106 connects to the target platform, the server 104 corresponding to the target platform sends a first data packet carrying a target identifier to the camera 106. If a return data packet carrying the target identifier is received through the first port, the intercom mode matched by camera 106 is determined to be the first intercom mode. If no return data packet carrying the target identifier is received on the first port, a second data packet carrying target address information is sent to camera 106. If a return data packet based on the target address information is received from camera 106, the intercom mode matched by camera 106 is determined to be the second intercom mode. After determining the intercom mode corresponding to camera 106, server 104 can send audio data packaged with the target communication protocol to camera 106 based on the intercom mode matched to camera 106, so that camera 106 can receive the audio data and play the corresponding voice. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0059] In one exemplary embodiment, such as Figure 2 As shown, a voice intercom method for a camera is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 202 to 208. Wherein:
[0060] Step 202: When the camera is connected to the target platform, the server corresponding to the target platform sends a first data packet carrying the target identifier to the camera.
[0061] The target platform is used for unified management of multiple cameras. It enables functions such as live streaming, cloud recording, and two-way communication between the cameras. For cameras connected to the network, communication (e.g., data transmission) typically requires knowing the camera's network address, such as its IP address and port number. Therefore, when a camera connects to the target platform, the server can attempt to send data packets to the network. If a response data packet from the camera is received, parsing the response packet reveals the camera's address. Furthermore, receiving a response data packet from the camera indicates the transmission protocol used by the server to send data to the camera, and the camera also supports this protocol. Therefore, audio data can then be sent to the previously parsed address using this transmission protocol, enabling two-way voice communication between the target platform and the camera.
[0062] The server supports two intercom modes: a first intercom mode and a second intercom mode. The first intercom mode is based on the target identifier, while the second intercom mode is based on the target address information.
[0063] In the first intercom mode, the server sends a first data packet carrying a target identifier to the camera. If the camera supports the first intercom mode, it will send a response data packet to the server after receiving the first data packet. This response data packet carries the target identifier and the camera's address information. Therefore, upon receiving this response data packet, the server can determine the camera's address information. The server can then communicate with the camera using the first intercom mode and the camera's address information.
[0064] In the second intercom mode, the server sends a second data packet carrying the target address information to the camera. Since the target address information is the server's address information, the camera, upon receiving the target address information, can send a data packet to the server based on the target address information, thereby establishing a communication connection with the server. The server can then communicate with the camera through this established connection.
[0065] For example, when a camera is connected to the target platform, the server can first attempt to establish a connection with the camera in a first intercom mode. Specifically, the server can send a first data packet carrying a target identifier to the camera. Optionally, the target identifier can be an SSRC (Synchronization Source Identifier). If the camera supports the first intercom mode, the communication data packet between the camera and the server will carry the SSRC. In this way, the server can determine the camera's address information through the SSRC carried in the data packet, thereby enabling intercom communication with the camera.
[0066] Step 204: If a return data packet carrying a target identifier is received through the first port, then the intercom mode matched by the camera is determined to be the first intercom mode.
[0067] In the first intercom mode, communication with the camera is achieved through the first port.
[0068] In some embodiments, the server may include an intercom server. When a camera connects to the target platform, the server generates a target identifier and sends it to both the intercom server and the camera. Optionally, the server may send an Invite request to the camera via the SIP protocol. This Invite request carries the target identifier. After receiving the Invite request, the camera sends a response data packet to the intercom server according to the given target identifier. If the intercom server receives a response data packet carrying the target identifier on its first port, and this target identifier is the aforementioned generated identifier, the intercom server can then communicate with the camera based on the address information parsed from the response data packet.
[0069] Step 206: If no return data packet carrying the target identifier is received at the first port, a second data packet carrying the target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode.
[0070] The second intercom mode communicates with the camera via a second port.
[0071] Since not all cameras support the RTP transmission protocol or the target identifier-based intercom mode, a second intercom mode needs to be attempted to communicate with these cameras. Because the second intercom mode uses a common network transmission protocol like TCP, cameras that cannot use the first intercom mode can usually communicate using the second intercom mode.
[0072] In some embodiments, the server can send an Invite request to the camera via the SIP protocol, informing the camera of the intercom server's address information (i.e., the target address information). Upon receiving the Invite request, the camera proactively sends a TCP connection establishment request to the intercom server. After receiving the camera's TCP connection establishment request, the intercom server allocates a dedicated intercom port to the camera. Thus, the intercom server can communicate with the camera through the allocated dedicated intercom port.
[0073] Step 208: Based on the intercom mode matched with the camera, send audio data packaged with the target communication protocol to the camera so that the camera receives the audio data and plays the corresponding voice.
[0074] After the above steps, the server has obtained the camera's address information and the intercom modes supported by the camera. The intercom modes supported by the camera can be stored on the server as the camera's tag information. Subsequently, the server can send audio data packaged with the target communication protocol to the camera based on the intercom mode matched by the camera, so that the camera can receive the audio data and play the corresponding voice.
[0075] The target communication protocol can be the RTP communication protocol. For both intercom modes, the target communication protocol is used to package the audio data. This standardizes the intercom rules, allowing different devices to communicate in a standardized way, thus simplifying voice intercom compatibility issues.
[0076] In the aforementioned voice intercom method for cameras, when a camera connects to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera. If a return data packet carrying a target identifier is received through the first port, the intercom mode matched by the camera is determined to be the first intercom mode. If no return data packet carrying a target identifier is received at the first port, a second data packet carrying target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. In this way, after the camera connects to the target platform, it attempts to determine which of the two intercom modes the camera supports. Compared to obtaining detailed information about the camera for complex adaptation and configuration of each camera, this simplifies the configuration process required for the camera to connect to the target platform and enables adaptation of intercom modes for cameras with different sound fields from different manufacturers. This improves platform compatibility, reduces the possibility of manual intervention and configuration errors, improves the reliability and stability of the target platform, and reduces operation and maintenance costs.
[0077] In one exemplary embodiment, such as Figure 3 As shown, when a camera is connected to a target platform, the steps for the server corresponding to the target platform to send first data carrying a target identifier to the camera include steps 302 to 306. Wherein:
[0078] Step 302: When the camera is connected to the target platform, the server generates a target identifier.
[0079] Step 304: Create a session establishment request packet carrying the target identifier.
[0080] Step 306: Send a session establishment request data packet to the camera via the SIP protocol.
[0081] The target identifier can be a synchronization source identifier. The server can then identify whether a corresponding target identifier exists in the received data packets. If it does, the server can determine that the camera supports the first intercom mode and obtain the camera's address information by parsing the received data packets.
[0082] SIP (Session Initialization Protocol) is a network communication protocol and the foundation of modern interactive communication on the Internet. A session establishment request packet can be, for example, the Invite request packet mentioned above.
[0083] For example, when a camera connects to the target platform, the server generates a target identifier and sends the target identifier to the camera via a session establishment request packet, and also sends the target identifier to the intercom server. Then, the intercom server parses the identifier carried in the received data packet on the first port to determine if it is the target identifier.
[0084] In this embodiment, by generating a target identifier and creating a session establishment request data packet carrying the target identifier, and then sending the session establishment request data packet to the camera via the SIP protocol, the intercom service in the server can determine the camera's address information and supported intercom modes based on the target identifier. This allows the camera to determine its supported voice intercom modes through matching, simplifying the device configuration process.
[0085] In some embodiments, such as Figure 4 As shown, the steps for sending a second data packet carrying target address information to the camera include:
[0086] Step S402: Create a session establishment request data packet carrying the target address information.
[0087] Step S404: Send a session establishment request data packet to the camera via the SIP protocol.
[0088] Step S406: Receive a TCP connection establishment request data packet from the camera carrying the target address information.
[0089] The target address information can be the server's address information. Optionally, the target address information can be the address information of the intercom server. The second data packet is sent based on the TCP protocol. The second intercom mode can be a TCP connection-based intercom mode. To establish intercom communication between the camera and the server, the server can send a session request packet to the camera via the SIP protocol, wherein the session request packet carries the target address information. After receiving the session request packet, the camera can send a TCP connection request to the intercom server. After receiving the TCP connection request, the intercom server allocates a separate intercom port to communicate with the camera.
[0090] In this embodiment, a session establishment request data packet carrying target address information is created and sent to the camera via the SIP protocol. After the camera receives the session establishment request data packet, the server receives the TCP connection establishment request sent by the camera. In this way, the intercom server can obtain the camera's address information and the information that the camera supports the second intercom mode. This makes it easier for the intercom server to communicate with the camera directly based on the camera's address information and the intercom rules corresponding to the second intercom mode. This simplifies the complex configuration process of the camera, and the camera configuration is completed by selecting the intercom mode that matches the camera from the two intercom modes.
[0091] In some embodiments, after obtaining the camera's address based on a TCP connection establishment request packet, the server allocates a separate target intercom port to the camera to send audio data to the camera's address through the target intercom port. Through these steps, the server can establish reliable intercom communication with the camera, and communicating with the camera through a separate target intercom port also improves the communication efficiency and quality of the camera intercom.
[0092] As mentioned earlier, the server packages the audio data sent to the camera using the target communication protocol. The target communication protocol could be, for example, RTP. The packaging rules can be that the audio payload data corresponds to an integer number of audio encoded frames, and the audio bitstream conforms to the target encoding standard. For example, the duration of the audio payload data could be between 20 and 180 milliseconds. Alternatively, the duration could be 100 milliseconds or 80 milliseconds. By standardizing the format of the transmitted audio data, the communication between the camera and the server can conform to the standard, improving the reliability of the server's voice communication.
[0093] In some embodiments, since different manufacturers configure cameras with different rules for parsing audio data, some format processing of the audio data is required. When matching the intercom mode of the camera, the server can also know the manufacturer of the camera or the device type of the camera. For some device types, the corresponding data packets that can be parsed may be in a first format. For other device types, the corresponding data packets that can be parsed may be in a second format. Since the target communication protocol is the RTP protocol, the first format may include an RTP header. The second format may not include an RTP header. Optionally, the first format also needs to add two bytes of RTP packet length, while the second format does not need to add two bytes of RTP packet length. For example, before sending audio data to the camera, the server first obtains the device type of the camera; determines the data format of the audio data based on the device type; and sends audio data in the data format packaged by the RTP protocol to the camera, so that the camera can play the corresponding voice when it receives the audio data in the data format.
[0094] The data format is selected from either a first format or a second format. This ensures that the audio data is compatible with the camera's data parsing rules, allowing the camera to parse the audio data sent by the server and play the corresponding voice correctly.
[0095] In an exemplary embodiment, the target platform corresponding to the camera's voice intercom method consists of the following parts.
[0096] (1) Video monitoring terminal: responsible for collecting audio data from the client and encrypting it before sending it to the voice intercom gateway.
[0097] (2) Voice intercom gateway: Receives and decrypts audio data, and sends the decrypted audio data to the intercom service.
[0098] (3) Intercom Service: Responsible for processing and distributing audio data. It receives RTP packets from cameras, parses SSRC to find the corresponding intercom device, or receives TCP connections. It finds the matching camera through a fixed port and parses the camera's IP and port information. Based on the parsed camera's IP and port information, it sends audio data to the camera.
[0099] (4) Camera: A third-party manufacturer provides a camera that supports different intercom functions and is responsible for receiving and playing audio data.
[0100] To enable voice communication with the camera, as shown in Figure 5, when the camera connects to the target platform, the server generates a target identifier; creates a session establishment request data packet carrying the target identifier; and sends the session establishment request data packet to the camera via the SIP protocol. If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port. For example, the video surveillance platform initiates an Invite request to the camera via the SIP protocol, carrying a predetermined SSRC. After receiving the Invite request, the camera sends an RTP packet to the intercom service according to the given SSRC. After receiving the RTP packet on its fixed intercom port, the intercom service parses the SSRC in the RTP header. If the parsed SSRC matches a device that has started intercom, it is determined that the camera is matched with the first intercom mode.
[0101] If no return data packet carrying the target identifier is received on the first port, a session establishment request data packet carrying the target address information is created (see Figure 6). The session establishment request data packet is sent to the camera via the SIP protocol. If a TCP connection establishment request data packet carrying the target address information is received from the camera, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port. For example, the video surveillance platform initiates an Invite request to the camera via the SIP protocol, informing the camera of the receiving IP and port of the intercom service (e.g., 192.168.1.100:5000). After receiving the Invite request, the camera actively sends a TCP connection establishment request to the intercom service. After receiving the TCP connection establishment request from the camera, the intercom service allocates a separate intercom port (e.g., 6000) to the camera. After receiving the TCP connection establishment on the specific port, the intercom service resolves the camera's outgoing IP and port. This allows it to determine that the intercom mode matched by the camera is the second intercom mode.
[0102] After determining the intercom mode matched by the camera, the server, i.e., the intercom service, sends audio data packaged with the target communication protocol to the camera based on the matched intercom mode, enabling the camera to receive the audio data and play the corresponding voice. Regardless of the intercom mode used, the intercom service packages the audio data into RTP packets and sends them to the camera. The specific processing is as follows: The audio data collected by the front-end device (the audio receiving device connected to the video surveillance platform) is transmitted to the video surveillance platform. The video surveillance platform encrypts the collected audio data and sends it to the voice intercom gateway. The voice intercom gateway receives the encrypted audio data, decrypts it, and sends the decrypted data to the intercom service. The intercom service packages the audio data according to the pre-configured RTP packaging rules and the RTP protocol. The audio payload data should be an integer number of audio encoded frames, with a duration between 20ms and 180ms. The packaged RTP packets are transmitted to the target camera over the network.
[0103] Furthermore, the way audio data is processed differs depending on the type of camera. For the first type of camera, an RTP header and a 2-byte RTP packet length are required. For the second type of camera, an RTP header and the 2-byte length are not needed for normal audio transmission. When sending audio data, the intercom service on the server dynamically adjusts the data format according to the camera type to ensure that the audio data can be correctly received and played by the camera.
[0104] Through the above steps, the intercom function can be dynamically expanded based on the device type. For different types of cameras, the processing logic of the intercom service can be dynamically adjusted according to their supported intercom functions and audio data reception and processing methods, ensuring that the system can adapt to camera devices from more manufacturers. This solves the compatibility issues existing in the current system, simplifies the system adaptation and configuration process, improves system stability and user experience, and provides a highly efficient and reliable intelligent voice intercom solution for video surveillance systems.
[0105] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0106] Based on the same inventive concept, this application also provides a camera voice intercom device 700 for implementing the camera voice intercom method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more camera voice intercom device embodiments provided below can be found in the limitations of the camera voice intercom method described above, and will not be repeated here.
[0107] In an exemplary embodiment, as shown in FIG7, a camera-based voice intercom device is provided, comprising: a first determining module 701, a second determining module 702, and an intercom module 703, wherein:
[0108] The first determining module 701 is used to send a first data packet carrying a target identifier to the camera when the camera is connected to the target platform; if a return data packet carrying a target identifier is received through the first port, the intercom mode matched by the camera is determined to be the first intercom mode, and the first intercom mode communicates with the camera through the first port.
[0109] The second determining module 702 is used to send a second data packet carrying target address information to the camera if no return data packet carrying the target identifier is received at the first port, and to determine the intercom mode matched by the camera as the second intercom mode if a return data packet carrying the target address information is received from the camera. The second intercom mode communicates with the camera through the second port.
[0110] The intercom module 703 is used to send audio data packaged with the target communication protocol to the camera in an intercom mode based on camera matching, so that the camera receives the audio data and plays the corresponding voice.
[0111] In some embodiments, the first determining module is further configured to: when the camera is connected to the target platform, the server generates a target identifier; create a session establishment request data packet carrying the target identifier; and send the session establishment request data packet to the camera via the SIP protocol.
[0112] In some embodiments, the second determining module 702 is further configured to create a session establishment request data packet carrying target address information; send a session establishment request data packet to the camera via the SIP protocol; and receive a TCP connection establishment request data packet from the camera carrying target address information.
[0113] In some embodiments, the second determining module 702 is further configured to obtain the address of the camera based on the TCP connection establishment request data packet; and allocate a separate target intercom port to the camera to send audio data to the address of the camera through the target intercom port.
[0114] In some embodiments, the camera's voice intercom device further includes a format configuration module, which is used to obtain the camera's device type; determine the data format of the audio data based on the device type; and send audio data in a data format packaged via the RTP protocol to the camera, so that the camera can play the voice corresponding to the audio data when it receives the audio data in the data format.
[0115] The various modules in the aforementioned camera's voice intercom device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0116] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a camera-based voice intercom method.
[0117] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0118] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0119] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0120] If a return data packet carrying a target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port.
[0121] If no return data packet carrying the target identifier is received at the first port, a second data packet carrying the target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0122] Based on the camera-matching intercom mode, audio data packaged with the target communication protocol is sent to the camera so that the camera can receive the audio data and play the corresponding voice.
[0123] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0124] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0125] If a return data packet carrying a target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port.
[0126] If no return data packet carrying the target identifier is received at the first port, a second data packet carrying the target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0127] Based on the camera-matching intercom mode, audio data packaged with the target communication protocol is sent to the camera so that the camera can receive the audio data and play the corresponding voice.
[0128] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0129] When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera.
[0130] If a return data packet carrying a target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port.
[0131] If no return data packet carrying the target identifier is received at the first port, a second data packet carrying the target address information is sent to the camera. If a return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port.
[0132] Based on the camera-matching intercom mode, audio data packaged with the target communication protocol is sent to the camera so that the camera can receive the audio data and play the corresponding voice.
[0133] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0136] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A voice intercom method for a camera, characterized in that, The method includes: When a camera is connected to a target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera. If a return data packet carrying the target identifier is received through the first port, it is determined that the intercom mode matched by the camera is the first intercom mode, and the first intercom mode communicates with the camera through the first port; If the return data packet carrying the target identifier is not received at the first port, a second data packet carrying the target address information is sent to the camera. If the return data packet sent by the camera based on the target address information is received, the intercom mode matched by the camera is determined to be the second intercom mode. The second intercom mode communicates with the camera through the second port. Based on the intercom mode matched by the camera, audio data packaged with the target communication protocol is sent to the camera so that the camera receives the audio data and plays the voice corresponding to the audio data.
2. The method according to claim 1, characterized in that, When the camera connects to the target platform, the server corresponding to the target platform sends a first data packet carrying a target identifier to the camera, including: When the camera is connected to the target platform, the server generates the target identifier; Create a session establishment request packet carrying the target identifier; The session establishment request data packet is sent to the camera via the SIP protocol.
3. The method according to claim 2, characterized in that, The target identifier is the synchronization source identifier.
4. The method according to claim 1, characterized in that, Sending the second data packet carrying the target address information to the camera includes: Create a session establishment request data packet carrying the target address information; The session establishment request data packet is sent to the camera via the SIP protocol; The receipt of the return data packet sent by the camera based on the target address information includes: The camera received a TCP connection establishment request data packet carrying the target address information.
5. The method according to claim 4, characterized in that, After receiving the TCP connection establishment request data packet carrying the target address information from the camera, the method further includes: The address of the camera is obtained based on the TCP connection establishment request data packet; A separate target intercom port is assigned to the camera to send the audio data to the address of the camera via the target intercom port.
6. The method according to claim 1, characterized in that, The audio payload data corresponding to the audio data consists of an integer number of audio encoded frames, and the audio bitstream corresponding to the audio data conforms to the target encoding standard.
7. The method according to claim 1, characterized in that, Sending audio data packaged using a target communication protocol to the camera, so that the camera receives the audio data and plays the corresponding speech, includes: Obtain the device type of the camera; The data format of the audio data is determined based on the device type; The system sends audio data in the specified data format, packaged via the RTP protocol, to the camera so that the camera can play the corresponding audio data when it receives the audio data.
8. A voice intercom device for a camera, characterized in that, The device includes: The first determining module is used to send a first data packet carrying a target identifier to the camera when the camera is connected to the target platform; if a return data packet carrying the target identifier is received through the first port, the intercom mode matched by the camera is determined to be the first intercom mode, and the first intercom mode communicates with the camera through the first port. The second determining module is configured to send a second data packet carrying target address information to the camera if the return data packet carrying the target identifier is not received at the first port, and if the return data packet sent by the camera based on the target address information is received, determine that the intercom mode matched by the camera is the second intercom mode, and the second intercom mode communicates with the camera through the second port. The intercom module is used to send audio data packaged with a target communication protocol to the camera based on the intercom mode matched by the camera, so that the camera receives the audio data and plays the voice corresponding to the audio data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video transmission method and device, electronic equipment and computer readable storage medium
CN113873342A
Elevator monitoring camera access call center system capable of achieving two-way conversation
CN114584535A