Communication system

The integration of intercom and surveillance camera systems with SIP and ONVIF protocols addresses the lack of combined functionality, providing enhanced recording and event detection capabilities.

JP7851703B2Active Publication Date: 2026-04-27TOA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOA CORP
Filing Date
2021-09-28
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing communication systems lack integration and enhanced functionality between intercom and surveillance camera systems, limiting their combined utility and value.

Method used

A communication system that integrates call terminals with surveillance cameras, allowing for simultaneous voice and video recording, event detection, and coordinated response across multiple networks, utilizing protocols like SIP and ONVIF for seamless operation.

Benefits of technology

Enables a communication system with higher added value through integrated voice and video recording, event detection, and coordinated responses, enhancing security and communication capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851703000001
    Figure 0007851703000001
  • Figure 0007851703000002
    Figure 0007851703000002
  • Figure 0007851703000003
    Figure 0007851703000003
Patent Text Reader

Abstract

To provide a communication system with a higher added value by fusing an intercom system or a monitor camera, for example.SOLUTION: A communication system 1 includes: a call terminal belonging to an intercom system (call system 10) for calling with another call terminal by using a call protocol corresponding to a voice and a video; a first recorder (recorder 30) for receiving a video transferred by using a monitor camera protocol corresponding to the video and storing the video; and a second recorder (recorder 42) for receiving a voice and a video transferred by using the call protocol and storing the voice and the video. The call terminal has a camera and transfers a video taken by the camera to the first recorder so that the first recorder records the taken image for crime prevention. The call terminal transfers a call voice and the image taken by the camera to the second recorder by using the call protocol at the time of calling so that the second recorder records the call voice and the taken image for a call log.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a communication system that can integrally operate in-building communication using an intercom or an IP speaker, public address broadcasting, video surveillance using a surveillance camera system, video recording, etc.

Background Art

[0002] An intercom system is a system mainly for in-building communication. Typically, an intercom system includes one or more call terminals and public address devices such as one or more speakers. Each call terminal has a microphone and a speaker and is interconnected via a network. Each call terminal conducts in-building communication by transmitting the input voice to a call terminal of another destination according to a predetermined protocol such as the SIP protocol. Also, each call terminal conducts a public address method from the public address device by transmitting the input voice to the public address device.

[0003] A surveillance camera system is a system mainly used for security purposes. Typically, a surveillance camera system includes one or more surveillance cameras and a recorder to which each surveillance camera is connected. Each surveillance camera includes a camera installed at a predetermined location such as on the ceiling or wall surface, or a camera installed on a moving body such as a drone or a vehicle. The recorder can record by receiving the captured video of the surveillance camera and storing it in a built-in or external recording medium, or can perform live viewing by outputting the captured video to an external display in real time.

[0004] An example of an intercom system is disclosed in Patent Document 1. An example of a surveillance camera system is disclosed in Patent Document 2.

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention aims to provide a communication system with higher added value by integrating intercom systems and surveillance camera systems, among other things. [Means for solving the problem]

[0006] The communication system belongs to the intercom system and comprises a call terminal that makes calls with other call terminals using a call protocol that supports voice and video, a first recording device that receives and stores video transmitted using a surveillance camera protocol that supports video, and a second recording device that receives and stores voice and video transmitted using the call protocol. The call terminal is equipped with a camera, and the call terminal transmits the video captured by the camera to the first recording device, so that the first recording device records the captured video for security purposes. During a call, the call terminal transmits the video captured by the camera along with the voice call using the call protocol to the second recording device, so that the second recording device records the voice call and the captured video as a call record. [Effects of the Invention]

[0007] This invention makes it possible to provide a communication system with higher added value. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 shows a communication system according to an embodiment of the present invention. [Figure 2] Figure 2 shows an example of a communication system configuration including LAN and WAN. [Figure 3] Figure 3 shows an example of the configuration of a call terminal within a call system. [Figure 4] Figure 4 shows an example of different recording operations depending on the operating mode. [Figure 5] Figure 5 shows an example of different recording operations depending on the operating mode. [Figure 6] Figure 6 is a flowchart showing an example of event detection functionality between a call system and a surveillance camera system. [Figure 7] Figure 7 shows an example of a database that maps the locations of the call area of ​​a call system to the locations of the surveillance area of ​​a surveillance camera system. [Figure 8] Figure 8 shows an example of a database that maps the locations of the call area of ​​a call system to the locations of the surveillance area of ​​a surveillance camera system. [Figure 9] Figure 9 shows an example of the configuration of surveillance cameras within a surveillance camera system. [Figure 10] Figure 10 shows an example configuration of a communication system according to an embodiment of the present invention, which includes an AI processing unit. [Figure 11] Figure 11 shows an example of the configuration of IP speakers within a public address system. [Figure 12] Figure 12 is a flowchart showing an example of a recorded broadcasting function in a telephone communication system. [Figure 13] Figure 13 is a flowchart showing an example of a download broadcast function between a call system and a public address system. [Figure 14] Figure 14 is a flowchart illustrating an example of a simultaneous broadcasting function between a call system and a public address system. [Figure 15] Figure 15 shows an example configuration in which a paging gateway within a public address system is connected to an emergency signal source. [Figure 16] Figure 16 is a flowchart showing an example of the operation flow of a paging gateway and IP speaker in a coexistence function with emergency broadcasting. [Figure 17] Figure 17 is a flowchart showing an example of the operation flow of a paging gateway in a function that coexists with emergency broadcasts. [Figure 18] Figure 18 is a flowchart showing an example of the operation flow of an IP speaker in a coexistence function with emergency broadcasting.

Best Mode for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 shows a communication system 1 according to an embodiment. The communication system 1 includes a call system 10, a system manager 20, a recorder 30, a surveillance camera system 40, and a voice amplification system 50. The call system 10, the system manager 20, the recorder 30, the surveillance camera system 40, and the voice amplification system 50 are interconnected via a network.

[0010] The call system 10 includes one or more call terminals 11. The call terminal 11 is a terminal for making a voice call and has a function of making a voice call by transmitting and receiving voice data to and from other call terminals 11 via a network. Each call terminal 11 is assigned an identifier (e.g., an IP address) that uniquely identifies itself on the network. The call terminal 11 transmits and receives voice data and video data in accordance with a protocol for IP calls during a voice call. An example of a protocol for IP calls is SIP (Session Initiation Protocol) standardized by the IETF (Internet Engineering Task Force) and defined in RFC3261.

[0011] The system manager 20 is a control device for the call terminal 11 to perform transmission and reception processing of voice data and video data in accordance with a protocol for IP calls. The transmission and reception processing includes call control such as originating a call. For example, the system manager 20 is a SIP server that performs call control in accordance with SIP. The system manager 20 is assigned an identifier (e.g., an IP address) that uniquely identifies itself on the network.

[0012] The recorder 30 is a recording device equipped with a recording medium such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). Like the call terminal 11, the recorder 30 conforms to the IP call protocol and has the function of recording video data transmitted from the call terminal 11 in accordance with the IP call protocol. The recorder 30 is assigned an identifier (e.g., an IP address) that uniquely identifies it on the network.

[0013] The surveillance camera system 40 is a system for monitoring a predetermined area (surveillance area) for security purposes or the like, and includes one or more surveillance cameras 41 and a recorder 42. Each surveillance camera 41 has a function of generating video data from a video signal acquired by an image sensor such as a CCD or CMOS and transmitting it externally. The surveillance camera 14 streams and transmits video in accordance with the transmission protocol for the surveillance camera system. An example of the transmission protocol for the surveillance camera system is a protocol compliant with the ONVIF (Open Network Video Interface Forum) standard. Each surveillance camera 41 is assigned an identifier (e.g., IP address) that uniquely identifies itself on the network. The surveillance camera 41, the recorder 42 is a recording device equipped with a recording medium such as an HDD or SSD, and has a function of recording the video data transmitted from the connected surveillance camera 41. The recorder 42 receives and records video data in accordance with the transmission protocol for the surveillance camera system, similar to the surveillance camera 41. The recorder 42 is assigned unique identification information (e.g., IP address) on the network. The recorder 42 has a voice input function via a microphone, generates voice data from the voice input from the microphone, and has a function of transmitting it externally. The recorder 42 transmits voice data in accordance with the transmission protocol for the surveillance camera system. The surveillance camera system 40 may be integratedly managed by a VMS (Video Management System). When integratedly managed by a VMS, the surveillance camera 41 and the recorder 42 are controlled and managed by the VMS software installed on a PC (Personal Computer). When integratedly managed by a VMS, the recorder 42 may be a PC equipped with the VMS software and a recording medium such as an HDD or SSD. In this case, the voice data of the voice input from the microphone provided in the PC is transmitted externally.

[0014] The public address system 50 is a system installed in a designated area (public address area) for purposes such as evacuation guidance during disasters and playback of background music, and includes one or more IP speakers 51. Each IP speaker 51 is assigned unique network identification information (e.g., an IP address). The IP speakers 51 have the function of receiving, playing, and amplifying audio data in accordance with the same IP call protocol as the call terminal 11 and the same surveillance camera system transmission protocol as the recorder 42. The public address system 50 also includes a paging gateway 52 connected to the same network as the IP speakers 51. The paging gateway 52 is assigned unique network identification information (e.g., an IP address). The paging gateway 52 has the function of receiving audio data via the network and transferring the audio data to the IP speakers 51 via the network in accordance with the same IP call protocol as the call terminal 11 and the same surveillance camera system transmission protocol as the recorder 42. In other words, the paging gateway 52 can broadcast simultaneously by first receiving audio data intended for the IP speakers 51 and then transferring the audio data to the IP speakers 51. The paging gateway 52 may transfer audio data to the IP speaker 51 via multicast or broadcast. The IP speaker 51 receives and plays the digitized audio data via a wireless communication medium such as an Ethernet cable or cellular communication.

[0015] In this way, in the communication system 1, the call system 10, the system manager 20, the recorder 30, the surveillance camera system 40, and the public address system 50 are interconnected via a network, allowing voice calls between the call terminals 11 of the call system 10, voice broadcasts from the call terminals 11 of the call system 10 to the IP speakers 51 of the public address system 50, and voice broadcasts from the recorder 42 of the surveillance camera system 40 to the IP speakers 51 of the public address system.

[0016] As shown in Figure 2, the communication system 1 may be configured via multiple LANs (Local Area Networks) and a WAN (Wide Area Network) such as the Internet that interconnects them. The call system 10, system manager 20, recorder 30, surveillance camera system 40, and public address system 50 are each configured on the LAN and interconnected. In such a configuration, various forms of communication can be performed, such as voice calls between call terminals 11 between different call systems 10 across multiple LANs, voice broadcasts from a call terminal 11 of a call system 10 in one LAN to an IP speaker of a public address system 50 in another LAN, or voice broadcasts from a recorder 42 of a surveillance camera system 40 in one LAN to an IP speaker 51 of a public address system 50 in another LAN. When the call system 10 is configured on the LAN, the call terminal 11 is preferably connected to a router constituting the LAN via an Ethernet cable, transmitting voice data via the Ethernet cable and receiving PoE (Power over Ethernet) power via the Ethernet cable. If the surveillance camera system 40 is configured on a LAN, the surveillance cameras 41 and recorder 42 should be connected to the router that makes up the LAN via Ethernet cables, and video data should be transmitted and received via the Ethernet cables, as well as powered by PoE via the Ethernet cables. If the public address system 50 is configured on a LAN, the IP speakers 51 and paging gateway 52 should be connected to the router that makes up the LAN via Ethernet cables, and audio data should be received via the Ethernet cables, as well as powered by PoE via the Ethernet cables.

[0017] The following explanation uses SIP as the protocol for IP calls and the ONVIF protocol as the transmission protocol for surveillance camera systems.

[0018] As shown in Figure 3, the communication terminal 11 includes a microphone 110, a speaker 111, a camera 112, a display unit 113, and an input unit 114. The microphone 110 is an electroacoustic transducer whose sound pickup direction is directed outward from the housing of the communication terminal 11, and it picks up the voice of a speaker speaking to the communication terminal 11 and outputs an audio signal. The speaker 111 is an electroacoustic transducer whose sound emission direction is directed outward from the housing of the communication terminal 11, and it outputs the sound obtained by reproducing the audio data. The camera 112 is an imaging device that includes an image sensor such as a CCD image sensor or a CMOS image sensor, and its imaging direction is directed outward from the housing of the communication terminal 11, so as to capture the image of a speaker speaking to the communication terminal 11. The camera 112 captures images and generates and outputs still image or moving image data from the captured video signal. The display unit 113 is a display including a display element such as an LCD or organic EL, and displays images captured by the camera 112 of the terminal 11 or other communication terminals 11, as well as other images related to the GUI, etc. The input unit 114 is an input device that accepts user operations, such as buttons and switches. The input unit 114 and the display unit 113 may be integrated to form a touch panel. The call terminal 11 can take various forms, such as a type that sits on a desk or table, or a type that is wall-mounted or embedded in a wall. The call terminal 11 is, so to speak, a call device called a speakerphone, intercom, or intercom.

[0019] (Operating Mode) The communication terminal 11 operates in multiple operating modes. These operating modes include at least a standby mode, a communication mode, and a public address mode. The standby mode is the mode in which the communication terminal 11 operates after power-on, without making calls to other communication terminals 11, nor making public address broadcasts to the IP speaker 51 or paging gateway 52. ​​The standby mode includes a general standby mode and a monitoring standby mode. The general standby mode is a mode in which the microphone 110, speaker 111, and camera 112 are turned OFF, and neither audio nor video is transmitted externally. The monitoring standby mode is a mode in which at least video is transmitted to the surveillance camera system 40 for security purposes, and at least the camera 112 and optionally the microphone 110 are turned ON, while the speaker 111 is turned OFF. When in standby mode, the communication terminal 11 switches between the general standby mode and the monitoring standby mode manually in response to user operation from the input unit 114, or automatically according to predetermined conditions. For example, when the call terminal 11 is in standby mode, it operates in general standby mode by default and switches to monitoring standby mode in response to user operation from the input unit 114. For example, when the call terminal 11 is in standby mode, it operates in general standby mode by default and switches to monitoring standby mode in response to the arrival of a predetermined time according to a predetermined schedule.

[0020] The call terminal 11 switches to call mode or amplification mode manually in response to user operation from the input unit 114, or automatically according to predetermined conditions, while in standby mode. Call mode is a mode in which the call terminal 11 transmits and receives voice data with other call terminals 11, and makes calls between multiple call terminals 11. Amplification mode is a mode in which the call terminal 11 transmits voice data to the IP speaker 51 and amplifies the voice from the IP speaker 51. For example, the call terminal 11 transitions to call mode in response to an operation via the input unit 114 to select the destination of another call terminal 11 to call and make a call. For example, the call terminal 11 transitions to call mode in response to receiving a call from another call terminal 11. For example, the call terminal 11 transitions to call mode in response to a user operation via the input unit 114 to respond to a call from another call terminal 11. For example, the calling terminal 11 transitions to loudspeaker mode in response to an operation to select a destination, such as an IP speaker 51 or paging gateway 52, to be the recipient of the loudspeaker broadcast, via the input unit 114 and initiate a call. When the communication terminal 11 is in calling mode, it turns on the microphone 110, speaker 111, and camera 112 for voice calls (video calls). When the calling terminal 11 is in loudspeaker mode, it turns on the microphone 110 for loudspeaker broadcasting. When the calling terminal 11 is in calling mode, it exits calling mode and transitions to standby mode when it has finished communicating with the other party's calling terminal 11 and the call has ended. When the calling terminal 11 is in loudspeaker mode, it exits loudspeaker mode and transitions to standby mode when it has finished communicating with the IP speaker 51 or paging gateway 52.

[0021] (Call function) When in call mode, the call terminal 11 transmits audio data of voice input from the ON microphone 110 and video data of video captured by the camera 112 to the other call terminal 11 (the call terminal 11 of the other party) using SIP. Also, when in call mode, it plays back the audio data transmitted via SIP from the other call terminal 11 and outputs the sound from the speaker 111, and plays back the transmitted video data and displays it on the display unit 113. In parallel with transmitting to the call terminal 11 of the other party, the call terminal 11 also transmits the audio data of voice input from the microphone 110, the video data of video captured by the camera 112, and the audio and video data transmitted from the call terminal 11 of the other party to the call terminal 11 to the recorder 30 using SIP. The recorder 30 records the received audio and video data as a call record according to SIP. When the calling terminal 11 ends a call and exits call mode, it responds by ending the transmission and reception of audio and video data with the other party's calling terminal 11, and ending the transmission of audio and video data to the recorder 30. The recorder 30 starts recording when it begins receiving audio and video data from the communication terminal 11 that has entered call mode, and stops recording when the call mode ends and the reception of audio and video data from the communication terminal 11 stops, saving the video data generated during that time as a call record. The recorder 30 should generate one video file each time recording starts and ends, so that a separate video file is generated as a call record each time a call is made.

[0022] When in amplification mode, the communication terminal 11 transmits the audio data of the voice input from the ON microphone 110 to the destination IP speaker 51 or paging gateway 52 using SIP.

[0023] (Recording function) When in monitoring standby mode, the communication terminal 11 generates video data of the image captured by the ON camera 112 and audio data of the sound input from the microphone 110, and continuously transmits them to the recorder 42 using the ONVIF protocol. The recorder 42 records the received video data and audio data to a recording medium in accordance with the ONVIF protocol. The communication terminal 11 continues recording to the recorder 42 in monitoring standby mode even when transitioning to call mode or loudspeaker mode. The operation in monitoring standby mode will be explained with reference to Figures 4 and 5. While set to monitoring standby mode, the communication terminal 11 continuously transmits audio data and video data to the recorder 42 to perform recording for security purposes. The communication terminal 11 does not terminate the recording operation to the recorder 42 even when call mode starts while in monitoring standby mode. In call mode, the communication terminal 11 continues to transmit audio and video data to the recorder 42, while simultaneously transmitting and receiving audio and video data with another call terminal 11 (the other party in the call) for the purpose of making a call, and also transmits audio and video data to the recorder 30 for call recording. When call mode ends, the call terminal 11 stops transmitting audio and video data to the recorder 30 for call recording, but continues to transmit audio and video data to the recorder 42. Similarly, while in monitoring standby mode, the call terminal 11 does not stop recording to the recorder 42 even if the public address mode starts. In public address mode, the communication terminal 11 continues to transmit audio and video data to the recorder 42, while simultaneously transmitting audio data to the IP speaker 51 or paging gateway 52 for public address broadcasting. When public address mode ends, the call terminal 11 stops transmitting audio data to the IP speaker 51 or paging gateway 52, but continues to transmit audio and video data to the recorder 42.

[0024] (Event detection function) The communication system 10 and the surveillance camera system 40 mutually detect events and respond to such detections. Here, an event is an occurrence that can be inferred from audio or video and is presumed to have occurred in the location where the communication system 10 is installed (communication area) or the location where the surveillance camera system 40 is installed (surveillance area). An example of an event is a scream made by someone in the communication area or surveillance area, which can be inferred, for example, by analyzing the audio. Another example of an event is a suspicious person entering the communication area or surveillance area, which can be inferred, for example, by analyzing the video.

[0025] Figure 6 is a flowchart illustrating the event detection function between the communication system 10 and the surveillance camera system 40. In the surveillance camera system 40, the surveillance camera 41 captures images within the surveillance area and detects events within the surveillance area from the captured images (S10). In S10, the surveillance camera 41 analyzes the video signal input from the image sensor and detects abnormalities within the surveillance area. For example, the surveillance camera system 40 has a motion detection function that compares multiple images obtained by the surveillance camera 41 at predetermined time intervals (at a predetermined frame rate), identifies locations (pixel groups) where changes occur in the images, and detects moving objects such as people or vehicles at the locations corresponding to these pixel groups. For example, the surveillance camera system 40 has an object detection function that stores image data related to photographs of people or vehicles to be detected in advance, compares each of the multiple images obtained by the surveillance camera 41 at predetermined time intervals (at a predetermined frame rate) with the image data, and detects that the person or vehicle is in the surveillance area if there is a predetermined degree of match. The above detection function may be provided in the surveillance camera 41 or in the recorder 42. In this manner, when the surveillance camera system 40 detects an anomaly in the surveillance area from the video footage captured by the surveillance camera 41, it responds by sending a notification message to the communication system 10 to inform it of the anomaly (S11). In S11, the surveillance camera 41 or the recorder 42 sends a communication message to the communication terminal 11 via the network. Here, the surveillance camera system 40 sends a notification message to the communication system 10 located in the communication area corresponding to the surveillance area in which its system is installed. That is, in S11, the surveillance camera system 40 identifies the communication system 10 having a communication area in the same location as or closest to the surveillance area where the anomaly was detected, and sends a notification message to that communication system 10.

[0026] Communication system 1 may include multiple communication systems 10, each having a different communication area, and surveillance camera systems 40, each having a different monitoring area. In this case, communication system 1 maintains a database containing communication area location information indicating the location of the communication area of ​​each communication system 10, and monitoring area location information indicating the location of the monitoring area of ​​each surveillance camera system 40. Figure 7 shows a database 70 as an example. Database 70 maintains the location of each communication system 10 among the multiple communication systems 10, and the location of each surveillance camera system 40 among the multiple surveillance camera systems 40. In the example shown in Figure 5, in database 70, communication system 10-1 is associated with location AAAA, communication system 10-2 is associated with location BBBB, communication system 10-3 is associated with location CCCC, surveillance camera system 40-1 is associated with location EEEE, surveillance camera system 40-2 is associated with location AAAA, and surveillance camera system 40-3 is associated with location BBBB. This indicates that each system is installed at its associated location. In the database 70, information for each location may be manually entered by a user, for example, by the administrator of the communication system 10 or the surveillance camera system 40, or it may be automatically entered by the system. For example, the communication system 10 and the surveillance camera system 40 are equipped with a positioning system using GNSS (Global Navigation Satellite System) such as GPS (Global Positioning System), and the database 70 is automatically generated by referring to the location information obtained by the positioning system, which includes longitude and latitude. For example, the communication system 10 and the surveillance camera system 40 are equipped with an indoor positioning system using beacon signals such as Bluetooth, RFID (Radio Frequency Identification), ultrasound, geomagnetic fields, UWB (Ultra Wide Band) signals, etc., and the database 70 is automatically generated by referring to the indoor location information obtained by the indoor positioning system, which includes the location information obtained by the system.For example, the call system 10 receives user input of location information from an input unit 114 provided in the call terminal 11, and the received location information is entered into the database 70. For example, the surveillance camera system 40 receives user input of location information from an input means (user interface such as a button, switch, or keyboard) provided in the surveillance camera 41 or recorder 42, and the received location information is entered into the database 70. The call system 10 and the surveillance camera system 40 can be judged to be installed in the same location if their location information perfectly matches, and the call system 10 and the surveillance camera system 40 can be judged to be installed in the same location if their location information is substantially the same and within a predetermined margin of error.

[0027] Figure 8 shows another example of the database 70. In the example shown in Figure 6, the database 70 holds the locations associated with each call terminal 11 in each call system 10 and the locations associated with each surveillance camera 41 in each surveillance camera system 40. In the example shown in Figure 6, in the database 70, call terminal 11-1 in call system 10-1 is associated with location aaaa, call terminal 11-2 is associated with location bbbb, call terminal 11-3 is associated with location cccc, surveillance camera 41-1 in surveillance camera system 40-1 is associated with location eeee, surveillance camera 41-2 is associated with location aaaa, and surveillance camera 41-3 is associated with location bbbb. This indicates that each call terminal 11 and each surveillance camera 41 are installed at their associated locations. In the database 70, information for each location may be manually entered by a user, for example, by the administrator of the communication terminal 11 or the surveillance camera 41, or it may be automatically entered by the system. For example, each communication terminal 11 and each surveillance camera 41 are equipped with a positioning system using GNSS such as GPS, and the location information is such as longitude and latitude acquired by the positioning system, and the database 70 is automatically generated by referring to this location information. For example, each communication terminal 11 and each surveillance camera 41 are equipped with an indoor positioning system using beacon signals such as Bluetooth, RFID, ultrasound, geomagnetic, UWB signals, etc., and the indoor location information is acquired by the indoor positioning system, and the database 70 is automatically generated by referring to this indoor location information. For example, user input of location information is received from the input unit 114 of the communication terminal 11, and the received location information is entered into the database 70. For example, user input of location information is received from the input means (user interface such as buttons, switches, keyboards, etc.) of the surveillance camera 41, and the received location information is entered into the database 70. If the location information of the communication terminal 11 and the surveillance camera 41 perfectly matches, it can be determined that they are installed in the same location. Similarly, if the location information of the communication terminal 11 and the surveillance camera 41 is associated with substantially the same location information within a predetermined margin of error, it can be determined that they are installed in the same location.

[0028] In S11, the surveillance camera system 40 refers to the database 70 to identify a communication system 10 that has the same or the closest communication area as its own surveillance area, and sends a notification message to that communication system 10. For example, the surveillance camera system 40 identifies a communication system 10 that substantially matches or is associated with location information that is closest to its own location, and sends a notification message to a communication terminal 11 within that communication system 10 (for example, if there are multiple communication terminals 11, any, some, or all of the communication terminals 11). For example, the surveillance camera system 40 identifies a surveillance camera 41 that detected an event in S10, identifies a communication terminal 11 that substantially matches or is associated with location information that is closest to the location information of the surveillance camera 41, and sends a notification message to the identified communication terminal 11.

[0029] In the call system 10, the call terminal 11 (S12) that receives a notification message responds to the notification message and performs an event response operation (S13). An example of an event response operation is to start saving the audio signal input from the microphone 110 or the video signal captured and acquired by the camera 112. In S13, if the microphone 110 was OFF, the call terminal 11 responds to the notification message by turning on the microphone 110 and starting to record the audio input from the microphone 110. In S13, if the camera 112 was OFF, the call terminal 11 responds to the notification message by turning on the camera 112 and starting to record the video captured by the camera 112. The call terminal 11 may be equipped with an internal recording medium such as an HDD or SSD, or it may accept an external recording medium such as an SD card or USB memory (flash drive), and in S13, it may save the recorded audio data or recorded video data to these recording media. In S13, the communication terminal 11, while in general standby mode, may respond to a notification message, switch to monitoring standby mode, turn on the microphone 110 and camera 112, and send the recorded audio data and video data to the recorder 42 to start saving on the recorder 42. In S13, the communication terminal 11 may perform recording and video recording in association with information related to the responded notification message (for example, the identifier of the monitoring camera 41 or recorder 42 that sent the notification message, and the date and time the notification message was received). Through S10 to S13, when the monitoring camera system 40 detects an event in a certain monitoring area, the communication system 10 in the communication area corresponding to that monitoring area can perform an action in response to the event. For example, when the intrusion of a suspicious person is detected by the monitoring camera 41 in a certain monitoring area, the communication terminal 11 located in the same place as or near the monitoring area can start recording and video recording, so that the situation of the suspicious person can be recorded and video recorded.

[0030] Conversely to S10-S14, the call system 10 monitors the call area using the microphone 110 and camera 112 provided on the call terminal 11. It detects events within the call area from the audio input from the microphone 110 and the video captured by the camera 112 (S20), and the surveillance camera system 40 responds to these events (S23). In S20, the call system 10 analyzes the audio signal input from the camera 110 of the call terminal 11 and detects abnormalities within the call area. For example, the call system 10 has an abnormality detection function that performs frequency analysis of the audio signal input from the microphone 110 at predetermined time intervals and detects abnormalities when a specific frequency pattern corresponding to a specific sound such as a scream occurs, or it monitors the level (amplitude) of the audio signal input from the microphone 110 at predetermined time intervals and detects abnormalities when a level above a predetermined level occurs. Alternatively, the call system 10 has an abnormality detection function that analyzes the video signal input from the image sensor of the camera 112 and detects abnormalities within the call area. This is similar to the motion detection function, object detection function, etc., of the surveillance camera system 40 described above. These detection functions may be provided in each call terminal 11. In this way, when the call system 10 detects an abnormality in the call area using the microphone 110 and camera 112 provided in the call terminal 11, it responds by sending a notification message to the surveillance camera system 40 to inform them of the abnormality (S21). In S21, the call terminal 11 that detected the abnormality sends a call message via the network to the surveillance camera 41 or recorder 42. Here, the call system 10 sends a notification message to the surveillance camera system 40 located in the surveillance area corresponding to the call area in which its system is installed. That is, in S21, similar to what the surveillance camera system 40 does in S11, the call system 10 refers to the database 70 to identify the surveillance camera system 40 that has a surveillance area in the same location as or the closest location to the call area in which the abnormality was detected, and sends a notification message to that surveillance camera system 40.As an example, the call system 10 identifies a surveillance camera system 40 that substantially matches or is associated with location information of its own location, and sends a notification message to a surveillance camera 41 (for example, one, some, or all of the surveillance cameras 41 if there are multiple surveillance cameras 41) or recorder 42 within that surveillance camera system 40. As an example, the call system 10 identifies a call terminal 11 that detected an event in S20, identifies a surveillance camera 41 that substantially matches or is associated with location information of the call terminal 11, and sends a notification message to the identified surveillance camera 41.

[0031] As shown in Figure 9, the surveillance camera 41 comprises a camera 410 and a PTZ mechanism 411. The camera 410 is an imaging device that includes an image sensor such as a CCD image sensor or a CMOS image sensor, and its imaging direction is oriented outwards from the housing of the surveillance camera 41, so as to capture part or all of the surveillance area where the surveillance camera 41 is installed. The camera 410 captures images and generates and outputs still image or moving image data from the resulting video signal. The PTZ mechanism 411 is connected to the camera 410 and includes an actuator such as a motor, which drives the camera 410 to pan, tilt, and zoom. The surveillance camera 41 is, for example, a camera that can be mounted on a wall or ceiling, and can be a box-type, burette-type (bullet-type), dome-type, or other type of camera.

[0032] In the surveillance camera system 40, the surveillance camera 41 or recorder 42 that receives a notification message (S22) responds to the notification message and performs an event response operation (S23). An example of an event response operation is to drive the PTZ mechanism 411 of the surveillance camera 41 to pan, tilt, and zoom the camera 410, capture images within the surveillance area, and save the video data to the recorder 42. In S23, if the camera 410 of the surveillance camera 41 that received the notification message is OFF, it responds to the notification message by turning the camera 410 ON and starting to capture images. More specifically, in S23, in response to the notification message, the surveillance camera 41 pans, tilts, and zooms in the direction of the communication terminal 11 that sent the notification message and takes images. In S23, if the camera 410 of the connected surveillance camera 41 is OFF, the recorder 42 that received the notification message responds to the notification message by instructing the surveillance camera 41 to turn the camera 410 ON and start capturing images. More specifically, in S23, in response to the notification message, the recorder 42 instructs the connected surveillance camera 41 to pan, tilt, and zoom in the direction of the call terminal 11 that sent the notification message and perform recording. In S23, the surveillance camera 41 may save the video data in association with information about the responded notification message (for example, the identifier of the call terminal 11 that sent the notification message, the date and time the notification message was received, etc.). S20 to S23 enable the call system 10 to detect an event in a call area, and the surveillance camera system 40 in the corresponding surveillance area to perform an action in response to that event. For example, when a scream is detected by the call terminal 11 in a call area, the surveillance camera 41 located in the same location as or near the call area can start recording and record the situation in which the scream occurred.

[0033] (Utilization of AI) The communication system 10 may utilize AI (Artificial Intelligence) when performing event detection using cameras 110 and 112 equipped on the communication terminal 11. As shown in Figure 10, the communication terminal 11 is equipped with an AI processing unit 115. The communication terminal 11 is also connected to a server 72 equipped with an AI processing unit 720 via a WAN such as the Internet. The AI ​​processing unit 115 is a processor that has a neural network and a pre-trained model generated in advance by machine learning (including deep learning). The server 72 is a web server accessible via the WAN and is equipped with an AI processing unit 720. The AI ​​processing unit 720 is a processor that has a neural network and a pre-trained model generated in advance by machine learning (including deep learning). The AI ​​processing unit 720 is more expensive but has higher functionality and processing power than the AI ​​processing unit 115. The AI ​​processing unit 115 is less expensive but has lower functionality and processing power than the AI ​​processing unit 720.

[0034] The AI ​​processing unit 115 includes a learning model 115a. The learning model 115a is a model that accepts audio data and video data as input and makes some kind of judgment from the audio data and video data. The learning model 115a is configured to make a judgment of the first event from new audio data and video data input by taking a large amount of audio data and video data samples related to the first event as training data and learning the features of the audio data and video data related to the first event. The AI ​​processing unit 720 includes a learning model 720a. The learning model 720a is a model that accepts audio data and video data as input and makes some kind of judgment from the audio data and video data. The learning model 720a is configured to make a judgment of the second event from new audio data and video data input by taking a large amount of audio data and video data samples related to the second event as training data and learning the features of the audio data and video data related to the second event. Here, unlike the first event, the second event requires higher computational power to determine. In other words, the process of determining the second event is more computationally intensive than the process of determining the first event. The determination process using the AI ​​processing unit 115 is a local process that does not require communication over the WAN, so it is highly responsive to the call terminal 11, meaning that the time to obtain the determination result is short and it has high real-time capabilities. On the other hand, the determination process using the AI ​​processing unit 720 can perform more advanced determinations, but it requires communication over the WAN, so it is less responsive to the call terminal 11, meaning that the time to obtain the determination result is short and it has low real-time capabilities.

[0035] Based on the characteristics of the AI ​​processing unit 115 and AI processing unit 720 described above, the call terminal 11 uses the AI ​​processing unit 115 and the AI ​​processing unit 720 as appropriate. For example, the call terminal 11 uses either the AI ​​processing unit 115 or the AI ​​processing unit 720 depending on the situation. For example, the call terminal 11 uses the AI ​​processing unit 115 and the AI ​​processing unit 720 in conjunction. When the call terminal 11 uses the local AI processing unit 115, it inputs voice data or video data to the AI ​​processing unit 115 and receives the judgment result from the learning model 115a. When the call terminal 11 uses the AI ​​processing unit 720 via the network, it transmits voice data or video data to the AI ​​processing unit 720 via the WAN and receives the judgment result from the learning model 720a via the WAN.

[0036] For example, the first event is an audio event such as a scream, and the judgment process for the first event is the process of determining whether an event such as a scream has occurred from the audio data. The learning model 115a is configured to take sample audio data of events such as screams as training data, learn the characteristics of screams, and then determine whether a scream has occurred from new audio data input. The second event is an audio event such as the appearance of a moving object, and the judgment process for the second event is the process of determining whether an event such as the appearance of a moving object has occurred from the video data. The learning model 720a is configured to take sample video data of events such as the appearance of a moving object as training data, learn the characteristics of moving objects, and then determine whether a moving object has occurred from new video data input. The communication terminal 11 can manually switch between using the AI ​​processing unit 115 to determine whether a scream has occurred from the audio data of the microphone 110's input audio, or using the AI ​​processing unit 720 to determine whether a moving object has occurred from the video data of the camera 112's captured video, in response to user input via the input unit 114. Furthermore, when the microphone 110 is ON, the communication terminal 11 may use the AI ​​processing unit 115 to determine the presence of a scream or the like from the input sound of the microphone 110, and when the camera 112 is ON, it may use the AI ​​processing unit 720 to determine the presence of a moving object or the like from the image captured by the camera 112. When both the microphone 110 and the camera 112 are ON, the communication terminal 11 may execute one or both of the determination process using the AI ​​processing unit 115 and the determination process using the AI ​​determination unit 720. Note that the first event being an audio event and the second event being an image event is merely an example and is not limited to this; conversely, the first event may be an image event and the second event may be an audio event.

[0037] For example, the determination of the first event is a first-stage determination related to sound or video, and the determination of the second event is a second-stage determination related to sound or video, which is a more computationally intensive determination than the first-stage determination. For example, the determination of the first event is a process to determine the occurrence of a moving object from video data, and the determination of the second event is a process to determine the occurrence of a moving object of a specific object, such as a predetermined person or object. In other words, the determination of the first event can be any determination of the occurrence of any moving object, regardless of what it is, but the determination of the second event is a determination of a specific person (e.g., a wanted criminal) or object (e.g., a car with a specific license plate) that has been predetermined to be determined. As an example, the communication terminal 11 uses the AI ​​processing unit 115 and the AI ​​processing unit 720 in stages. For example, the communication terminal 11 keeps the AI ​​processing unit 115 running at all times and inputs the video captured by the camera 112 to the AI ​​processing unit 115 at predetermined time intervals to perform the determination of the first event. Next, when the AI ​​processing unit 115 determines the first event, the communication terminal 11 responds by sending the video data in which the AI ​​processing unit 115 determined the first event to the server 72, and then performs the determination process for the second event. By performing the determination in this stepwise manner, for example, the AI ​​processing unit 115 can first determine the occurrence of some kind of moving object, and then the AI ​​processing unit 720 can subsequently determine whether that moving object is a specific person or object. This allows for more efficient determination than having the AI ​​processing unit 720 perform the determination process continuously.

[0038] (Sound amplification function) In loudspeaker mode, the call system 10 transmits audio data to the loudspeaker system 50, causing the IP speaker 51 to broadcast the audio. For example, each call terminal 11 receives user input from the input unit 114 to specify a destination IP speaker 51 or paging gateway 52, and transmits audio data to one or more specified IP speakers 51 or paging gateway 52. ​​If the call system 10 wants to broadcast the same audio content from the IP speaker 51 in the loudspeaker system 50, it is appropriate to transmit the audio data to the paging gateway 52. ​​The paging gateway 52 forwards the received audio data to all IP speakers 51 connected to the same network, causing all IP speakers 51 to play and output the audio data. As shown in Figure 11, the IP speaker 51 comprises a speaker 510 and a storage unit 510. The speaker 510 is an electroacoustic converter whose sound emission direction is directed outward from the housing of the IP speaker 51, and outputs the sound obtained from the playback of the audio data. The memory unit 511 is a storage medium that temporarily or permanently holds voice data received from the call system 10, and is a memory such as ROM or RAM, a storage device such as an HDD or SSD, or a combination thereof.

[0039] (Real-time broadcasting function) The communication system 10 has a real-time broadcasting function that broadcasts the input audio from the microphone 110 immediately and in real time when in amplification mode. In real-time broadcasting, the communication system 10 transmits the audio data of the audio input from the microphone 110 as an immediate audio stream to the designated IP speaker 51 or paging gateway 52. ​​For example, the communication terminal 11 generates audio data from the audio input from the microphone 110 and immediately transmits the audio stream to the IP speaker 51 or paging gateway 52 using a streaming protocol such as RTP (Realtime Transport Protocol). The IP speaker 51 immediately plays and outputs the audio streams it receives sequentially. The paging gateway 52 forwards the audio streams it receives sequentially to the IP speakers 51 on the same network.

[0040] (Recorded broadcast function) The call system 10 has a recording broadcast function that, in amplification mode, temporarily holds (records) the input audio from the microphone 110 in the call terminal 11, and broadcasts it after the user (speaker) has listened to it again. Figure 12 is a flowchart of an example of the recording broadcast function. When the recording broadcast function is enabled, the call terminal 11 generates audio data from the input audio from the microphone 110, but instead of transmitting it as an immediate audio stream, it temporarily saves the audio data to a recording medium (S40). The call terminal 11 may be equipped with an internal recording medium such as an HDD or SSD, or it may accept an external recording medium such as an SD card or USB memory (flash drive), and in S40, it may save the recorded audio data to these recording media. In S40, the call terminal 11 may automatically start recording when it detects that audio input from the microphone 110 has begun, such as when the speaker starts speaking into the microphone 110. In S40, the call terminal 11 may also accept a user operation from the input unit 114 to instruct the start of recording, and start recording accordingly. In S40, the call terminal 11 may automatically stop recording if the voice input from the microphone 110 is interrupted because the speaker stops speaking into the microphone 110 after recording has started. In S40, the call terminal 11 may also stop recording if it receives a user operation from the input unit 114 to stop recording after recording has started. In S40, the call terminal 11 generates a series of audio data from the audio input from the start to the end of recording and saves it to the storage medium. The call terminal 11 may retain the audio data saved to the storage medium in S40 for a certain period of time, and automatically erase the audio data from the recording medium after that period of time has elapsed. The call terminal 11 can play back the retained audio data within a certain period of time while the audio data is retained on the recording medium (S41). For example, in S41, the call terminal 11 automatically plays back the audio data within a certain period of time in response to the end of recording in S40. For example, in S41, the call terminal 11 receives user input from the input unit 114 instructing playback within a certain period of time, and in response, performs playback. In S41, the call terminal 11 outputs the played audio from the speaker 111.After playback in S41, the call terminal 11 receives a user operation from the input unit 114 to instruct broadcasting or to instruct a user operation to re-record. In response to a user operation to instruct broadcasting, the call terminal 11 broadcasts the audio data played back in S41, that is, it transmits the audio stream of the audio data held on the recording medium to the designated IP speaker 51 or paging gateway 52 (S42). On the other hand, in response to a user operation to instruct a re-recording, the call terminal 11 erases the audio data held on the recording medium and returns to S40 to record the speaker's voice again. With the above recording and playback function, the user (speaker) can play back what they have said, listen to it to check it, and then broadcast it. If they are not satisfied with the content after listening, they can re-record and broadcast it.

[0041] (Download broadcasting function) The call system 10, in loudspeaker mode, does not transmit an audio stream using a streaming protocol, but has a download broadcast function that downloads an audio file to an IP speaker 51 or paging gateway 52 and plays the downloaded audio file. Figure 13 is a flowchart of an example of the download broadcast function. When the download broadcast function is enabled in loudspeaker mode, the call terminal 11 in the call system 10 records the audio input via the microphone 110, that is, it encodes the audio to generate audio data and stores it in a storage medium (S50). In S50, the call terminal 11 may automatically start recording when audio input from the microphone 110 begins, such as when the speaker starts speaking into the microphone 110. In S50, the call terminal 11 may accept a user operation from the input unit 114 to instruct the start of recording and start recording accordingly. In S50, after recording has started, the call terminal 11 may automatically stop recording when audio input from the microphone 110 is interrupted, such as when the speaker stops speaking into the microphone 110. In S50, the call terminal 11 may, after recording has started, receive a user command from the input unit 114 to stop recording and terminate the recording accordingly. In S50, the call terminal 11 generates a series of audio data from the audio input from the start to the end of recording and saves it to the storage medium. In S50, the call terminal 11 may generate an audio file according to a predetermined file format such as MP3 and retain the audio file.

[0042] Next, the call terminal 11 downloads the audio data stored on the recording medium to the designated IP speaker 51 or paging gateway 52 (S51). In S51, the call terminal 11 may automatically start the download after the recording is completed in S50. In S51, the call terminal 11 may also accept a user operation from the input unit 114 to instruct the download after the recording is completed in S50, and start the download in response. In S51, the call terminal 11 transfers the audio data to the IP speaker 51 or paging gateway 52 using a file transfer protocol such as HTTP (Hypertext Transfer Protocol) or FTP (File Transfer Protocol).

[0043] The IP speaker 51 or paging gateway 52 downloads audio data and stores the downloaded audio data on a recording medium (S52). In S52, the IP speaker 51 stores the downloaded audio data in the storage unit 511. The paging gateway 52, like the IP speaker 51, is equipped with a storage medium including memory such as ROM or RAM, storage such as HDD or SSD, or a combination thereof, and in S50, it stores the downloaded audio data on the storage medium. When the download of the audio data is complete, the IP speaker 51 or paging gateway 52 sends a message to the call terminal 11 indicating this (download completion notification message) (S53).

[0044] With audio data stored in the IP speaker 51 or paging gateway 52, the call system 10 can instruct the public address system 50 to play and broadcast the stored audio data (S55-S57). The call terminal 11 receives a user operation from the input unit 114 to instruct the broadcasting of the built-in sound source, and in response, sends a request message (built-in sound source broadcast request message) to the designated IP speaker 51 or paging gateway 52 to instruct the broadcasting of the built-in sound source (S55). The IP speaker 51 receives the built-in sound source broadcast request message (S56), and in response, plays the audio data stored in the storage unit 511 and outputs it from the speaker 510 to broadcast to the public address area (S57). If there are multiple audio data stored in the storage unit 511, in S55 the call terminal 11 receives a user operation from the input unit 114 to select the audio data, and as a result, information indicating the selected audio data (audio data identification information) is included in the built-in sound source broadcast request message. In S57, the IP speaker 51 refers to the voice data identification information included in the built-in sound source broadcast request message, reads the voice data indicated by the voice data identification information from the storage unit 511, and plays it back. When the paging gateway 52 receives the built-in sound source broadcast request message (S56), it responds by transferring the voice data held in the storage medium to the IP speaker 51 on the same network, causing the IP speaker 51 to play the voice data (S57). If there are multiple voice data held in the storage medium, in S55 the call terminal 11 accepts a user operation to select voice data from the input unit 114, and as a result, information indicating the selected voice data (voice data identification information) is included in the built-in sound source broadcast request message. In S57, the paging gateway 52 refers to the voice data identification information included in the built-in sound source broadcast request message, reads the voice data indicated by the voice data identification information from the storage medium, and transfers it. In S57, the paging gateway 52 may transfer the voice data by multicast or broadcast.With the download broadcasting function described above, audio data is downloaded to the IP speaker 51 or paging gateway 52 before playback. Therefore, even if the communication environment is not suitable for streaming playback, such as when the network bandwidth is narrow, good public address broadcasting can be achieved without sound interruptions.

[0045] As a variation of S50-S57, the public address system 50 may automatically play the audio data without waiting for instructions from the call system 10 once the download of the audio data is complete (S52) (S57). In this variation, once the download of the audio data is complete (S52), the IP speaker 51 stores the audio data in its storage medium and sends a download completion notification message, or, without sending a download completion notification message, automatically plays the audio data stored in the storage medium to perform public address broadcasting (S57). Also in this variation, once the download of the audio data is complete (S52), the paging gateway 52 stores the audio data in its storage medium and sends a download completion notification message, or, without sending a download completion notification message, automatically transfers the audio data stored in the storage medium to the IP speaker 51, so that the audio data is played on the IP speaker 51 (S57).

[0046] (Simultaneous broadcasting function) The call system 10 has a simultaneous broadcasting function that, when in loudspeaker mode, broadcasts the same content from each IP speaker 51 simultaneously at the same time if the loudspeaker system 50 includes multiple IP speakers 51. Figure 14 is a flowchart showing an example of the simultaneous broadcasting function. When the simultaneous broadcasting function is enabled in loudspeaker mode, the call terminal 11 in the call system 10 records the audio input via the microphone 110, that is, it encodes the audio to generate audio data and stores it on a storage medium (S60). The recording operation in S60 is the same as the recording operation in S50. Next, the call terminal 11 downloads the audio data stored on the recording medium to the multiple IP speakers 51 in the loudspeaker system 50 (S61). In S61, the call terminal 11 may receive a user operation from the input unit 114 to select two or more IP speakers 51, and in response, transmit the audio data to the selected IP speakers 51. In S61, the call terminal 11 may receive a user operation from the input unit 114 instructing a mass broadcast and, in response, transmit audio data to all IP speakers 51 in the public address system 50. In S61, the call terminal 11 may automatically start downloading after recording is completed in S60. In S61, the call terminal 11 may receive a user operation from the input unit 114 instructing a download after recording is completed in S60 and, in response, start downloading. In S61, the call terminal 11 downloads the audio data by transferring it to the IP speakers 51 using a file transfer protocol such as HTTP or FTP.

[0047] Each IP speaker 51 downloads audio data and stores the downloaded audio data on a recording medium (S62). In S62, the IP speaker 51 stores the downloaded audio data in the storage unit 511. When the download of the audio data is complete, the IP speaker 51 sends a message to the call terminal 11 indicating this (download completion notification message) (S63). After the call system 10 starts transmitting the audio data (S61), it waits to receive download completion notification messages from all of the destination IP speakers 51 (S64, S65). When the call terminal 11 receives download completion notification messages from all of the IP speakers 51 (S65: Yes), the call system 10 automatically instructs the public address system 50 to play and broadcast the downloaded audio data (S66). In S66, the call terminal 11, in response to receiving download completion notification messages from all of the IP speakers 51 that downloaded the audio data in S61 (S65: Yes), automatically sends a broadcast instruction message to all of these IP speakers 51 instructing them to broadcast.

[0048] Each IP speaker 51 receives a broadcast instruction message (S67) and, in response, plays back the audio data downloaded and stored in the storage unit 511 and outputs it from the speaker 510 to broadcast to the public address area (S68). S60 to S68 ensure simultaneous broadcasting when multiple IP speakers 51 broadcast simultaneously within the public address area. In other words, if multiple IP speakers 51 were to individually receive, play back, and broadcast the audio data stream, depending on the network conditions between the call system 10 and the public address system 50 (the network conditions between the call terminal 11 and each IP speaker 51), there may be differences in the speed at which the IP speakers 51 receive the audio stream, potentially resulting in a time lag in the sound output between the IP speakers 51. Furthermore, even if the method involves saving audio data as a file in the call system 10 and downloading it to the public address system 50 before playback, similarly, depending on the network conditions between the call system 10 and the public address system 50 (the network conditions between the call terminal 11 and each IP speaker 51), there may be differences in the time it takes for the download to be completed among the IP speakers 51, potentially resulting in a time lag in the sound output between the IP speakers 51. In contrast, according to S60-S68, after the audio data download is completed on all IP speakers 51 to be broadcasted, the broadcasting of the downloaded audio data is instructed simultaneously, allowing all IP speakers 51 to start broadcasting at approximately the same time.

[0049] (Coexistence function with emergency broadcasting) The public address system 50 can be installed in office buildings, factories, shopping malls, and other commercial facilities. Depending on the size of these facilities, fire safety regulations may require the installation of emergency broadcasting equipment. Emergency broadcasting equipment is configured to amplify emergency broadcast audio with an amplifier and output it from multiple emergency broadcasting speakers installed on each floor of the facility. When the public address system 50 is installed in a facility where emergency broadcasting equipment is installed, IP speakers 51 are also installed on each floor, but care must be taken to ensure that they do not interfere with the emergency broadcast audio output from the emergency broadcasting speakers when the emergency broadcasting equipment is activated. In other words, according to fire safety regulations, when the emergency broadcasting equipment is activated, the emergency broadcast audio output from the emergency broadcasting speakers must be clearly distinguishable from other sounds and audible. Therefore, when the emergency broadcasting equipment is activated, a mechanism is required to prevent audio from being output from other speakers installed in the same location as the emergency broadcasting speakers. Conventionally, so-called business broadcasting equipment installed in the same location as emergency broadcasting equipment has used power cut relays that receive signals from the emergency broadcasting equipment or automatic fire alarm equipment and cut off the power supply. However, in the public address system 50, the IP speaker 51 may be connected to the router via an Ethernet cable and receive PoE power. In such a configuration, cutting off the power to the IP speaker 51 would require cutting off the power to the router, but other devices besides the IP speaker 51 may also be connected to the router, and it may not be desirable to cut off the power to the router itself. In light of these circumstances, the public address system 50 has a broadcast control function that can coexist with emergency broadcasting equipment without cutting off the power.

[0050] As shown in Figure 15, in the public address system 50, the paging gateway 52 is connected to an emergency signal source 80 independently of the path connecting to the telephone system 10, etc., via the network. The emergency signal source 80 is a signal source that emits a signal indicating the occurrence of an emergency, such as a fire, which activates the emergency broadcasting equipment installed in the facility where the public address system 50 is installed. The emergency signal source 80 is, for example, an automatic fire alarm system or an emergency broadcasting system. The emergency signal source 80 is configured to emit an emergency signal to the outside when it detects the occurrence of an emergency, such as a fire, and the paging gateway 52 receives this emergency signal.

[0051] Figure 16 is a flowchart showing an example of a broadcast control function for coexistence with emergency broadcasting equipment. The paging gateway 52 constantly monitors at predetermined time intervals whether or not it has received an emergency signal from the emergency signal source 80 (S80). If it receives an emergency signal (S80: Yes), it responds by sending a broadcast prohibition message (broadcast prohibition message) to all IP speakers 51 on the same network (S81). The broadcast prohibition message should be sent to the IP speakers 51 using multicast or broadcast.

[0052] Each IP speaker 51 normally operates with its operating mode set to broadcast-enabled mode by default (S82). When in broadcast-enabled mode, the IP speaker 51 can broadcast as needed, and for example, it plays and outputs audio data according to the real-time broadcasting function, recorded broadcasting function, download broadcasting function, and simultaneous broadcasting function described above. For example, when in broadcast-enabled mode, the IP speaker 51 plays and outputs audio streams that it receives sequentially according to the real-time broadcasting function and recorded broadcasting function. For example, when in broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to broadcast instructions received according to the download broadcasting function (S56~S57). For example, when in broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to broadcast instructions received according to the simultaneous broadcasting function (S67~S68). On the other hand, when the IP speaker 51 receives a broadcast prohibition message from the paging gateway 52 (S83), it responds by setting its operating mode to broadcast prohibition mode (S84). When the IP speaker 51 is in broadcast-prohibition mode, it cannot broadcast public address and does not play or output audio data. For example, when the IP speaker 51 is in broadcast-prohibition mode, even if it receives an audio stream according to the real-time broadcasting function or the recorded broadcasting function, it will not play the audio stream. For example, when the IP speaker 51 is in broadcast-prohibition mode, even if it receives a broadcast instruction according to the download broadcasting function (S56), it will not play the audio data. For example, when the IP speaker 51 is in broadcast-prohibition mode, even if it receives a broadcast instruction according to the simultaneous broadcasting function (S67), it will not play the audio data. By controlling it in this way, when the need arises to make an emergency broadcast, the operating mode of the IP speaker 51 can be switched to broadcast-prohibition mode, preventing interference with the emergency broadcast.

[0053] Figures 17 and 18 are flowcharts illustrating an example of a broadcast control function for coexistence with emergency broadcasting equipment. The paging gateway 52 constantly monitors at predetermined time intervals whether or not it has received an emergency signal from the emergency signal source 80 (S91), and, unless an emergency signal has been received, constantly sends a message indicating broadcast permission (broadcast permission message) to all IP speakers 51 within the same network (S91: No, S90). The broadcast permission message should be continuously sent to the IP speakers 51 at predetermined time intervals using multicast or broadcast. When the paging gateway 52 detects the reception of an emergency signal from the emergency signal source 80 (S91: Yes), it immediately and automatically stops transmitting the emergency signal in response (S92).

[0054] Each IP speaker 51 constantly monitors whether it has received a broadcast permission message from the paging gateway 52 at predetermined time intervals (S95), and operates in broadcast-ready mode as long as it has received a broadcast permission message (S96). When in broadcast-ready mode, the IP speaker 51 can broadcast as needed and, for example, plays and outputs audio data according to the real-time broadcasting function, recorded broadcasting function, download broadcasting function, and simultaneous broadcasting function described above. For example, when in broadcast-ready mode, the IP speaker 51 plays and outputs audio streams that it receives sequentially according to the real-time broadcasting function and recorded broadcasting function. For example, when in broadcast-ready mode, the IP speaker 51 plays and outputs audio data in response to broadcast instructions received according to the download broadcasting function (S56-S57). For example, when in broadcast-ready mode, the IP speaker 51 plays and outputs audio data in response to broadcast instructions received according to the simultaneous broadcasting function (S67-S68). On the other hand, if each IP speaker 51 stops receiving broadcast permission messages from the paging gateway 52 (S95: No), it immediately and automatically sets its operating mode to broadcast prohibition mode in response (S97). When in broadcast prohibition mode, the IP speaker 51 cannot broadcast public address and does not play or output audio data. For example, when in broadcast prohibition mode, even if the IP speaker 51 receives an audio stream according to the real-time broadcast function or the recorded broadcast function, it will not play the audio stream. For example, when in broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction according to the download broadcast function (S56), it will not play the audio data. For example, when in broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction according to the simultaneous broadcast function (S67), it will not play the audio data.

[0055] As a variation of S95-S97, the IP speaker 51 may not constantly confirm receipt of the broadcast permission message (S95), but only at the time of broadcasting. That is, for example, the IP speaker 51 may, in response to starting to receive an audio stream according to the real-time broadcasting function or the recorded broadcasting function, confirm receipt of the broadcast permission message (S95) before starting playback of the audio stream. In this case, the IP speaker 51 may buffer the audio stream and only start playback of the buffered audio stream if it has confirmed that it has received the broadcast permission message (S95:Yes). Alternatively, for example, the IP speaker 51 may, in response to receiving a broadcast instruction according to the download broadcasting function, confirm receipt of the broadcast permission message (S95) before starting playback of the audio data according to this broadcast instruction (S57). In this case, the IP speaker 51 may only start playback of the audio data if it has confirmed that it has received the broadcast permission message (S95:Yes) (S57). For example, the IP speaker 51, upon receiving a broadcast instruction in accordance with the broadcasting function, confirms receipt of a broadcast permission message (S95) before starting playback of audio data in accordance with this broadcast instruction (S68). In this case, the IP speaker 51 should only start playback of audio data (S68) if it has confirmed that it has received a broadcast permission message (S95: Yes).

[0056] As described above, each IP speaker 51 plays and broadcasts audio data in response to a broadcast permission message sent from the paging gateway 52 when the paging gateway 52 does not receive an emergency signal (when emergency broadcasting is not activated). If the reception of broadcast permission messages is interrupted, the speaker determines that it would interfere with emergency broadcasting and does not play or broadcast audio data. With the broadcast control shown in S80 to S84, if the broadcast prohibition message sent from the paging gateway 52 by S81 does not reach the IP speaker 51 due to packet loss or the like, there is a risk that the operating mode of the IP speaker 51 cannot be transitioned to broadcast prohibition mode. The broadcast control shown in S90 to S97 can more reliably prevent interference with emergency broadcasting.

[0057] (supplement) The telephone terminal 11 comprises a processing unit including a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), and a memory that holds computer programs that the processing unit can execute. Various functions of the telephone terminal 11 are realized by the processing unit executing the computer programs held in the memory. The computer programs held in the memory include program instructions that execute the various functions mentioned above, such as call functions, recording functions, event detection functions, AI-based judgment functions, public address functions, real-time broadcasting functions, recorded broadcasting functions, download broadcasting functions, and mass broadcasting functions.

[0058] The recorder 30 comprises a processing unit including a CPU and MPU, and memory that holds computer programs that the processing unit can execute. The various functions of the recorder 30 are realized by the processing unit executing the computer programs held in memory. The computer programs held in memory include program instructions that execute the various functions described above, such as the recording function.

[0059] The surveillance camera 41 comprises a processing unit including a CPU and MPU, and memory that holds computer programs that the processing unit can execute. The various functions of the surveillance camera 41 are realized by the processing unit executing the computer programs held in memory. The computer programs held in memory include program instructions that execute the various functions described above, such as event detection.

[0060] The recorder 42 comprises a processing unit including a CPU and MPU, and memory that holds computer programs that the processing unit can execute. The various functions of the recorder 42 are realized by the processing unit executing the computer programs held in memory. The computer programs held in memory include program instructions that execute the various functions described above, such as recording and event detection.

[0061] The IP speaker 51 comprises a processing unit including a CPU and MPU, and memory that holds computer programs that the processing unit can execute. The various functions of the IP speaker 51 are realized by the processing unit executing the computer programs held in memory. The computer programs held in memory include program instructions that execute the various functions mentioned above, such as public address function, real-time broadcasting function, recorded broadcasting function, download broadcasting function, mass broadcasting function, and emergency broadcasting function.

[0062] The paging gateway 52 comprises a processing unit including a CPU and MPU, and memory that holds computer programs that the processing unit can execute. The various functions of the paging gateway 52 are realized by the processing unit executing the computer programs held in memory. The computer programs held in memory include program instructions that execute the various functions mentioned above, such as the public address function, real-time broadcasting function, recorded broadcasting function, download broadcasting function, mass broadcasting function, and coexistence function with emergency broadcasting. [Explanation of symbols]

[0063] 1. Communication System 10. Calling System 20 System Manager 30 Recorders 40 Surveillance camera systems 50 Public address systems

Claims

1. An intercom system for on-premises communication comprising one or more communication terminals and one or more loudspeakers, wherein the communication terminals have a communication mode for mutual communication between the communication terminals and a loudspeaker mode for amplification from the communication terminals to the loudspeakers, A communication terminal that makes calls to other communication terminals using a call protocol that supports voice and video, A first recording device that receives and stores video transmitted using a surveillance camera protocol compatible with video, The system includes a second recording device that receives and stores audio and video transmitted using the aforementioned communication protocol, The aforementioned communication terminal is equipped with a camera, When the communication terminal is in a monitoring standby mode, which can be set while it is neither in the call mode nor the loudspeaker mode, it continuously transmits the video footage captured by the camera to the first recording device, so that the first recording device records the captured video footage for security purposes. The aforementioned call terminal transmits the call audio and the captured video from the camera to the second recording device using the call protocol when in call mode, and the second recording device records the call audio and captured video as a call record in this intercom system.

2. The intercom system according to claim 1, An intercom system in which the communication terminal transmits captured video to the first recording device and the first recording device continues to record the captured video for security purposes, even when transitioning from the monitoring standby mode to the communication mode or the loudspeaker mode.

3. A telephone terminal, It has a call mode for making calls with other communication terminals and a voice amplification mode for amplifying sound to a public address system. Equipped with a microphone, speaker, and camera, In the monitoring standby mode, which can be set during standby when neither the aforementioned call mode nor the aforementioned loudspeaker mode, the camera's captured video is continuously transmitted to the first recording device using the surveillance camera protocol to record for security purposes. In the aforementioned call mode, the system uses a call protocol to transmit the audio picked up by the microphone to the other call terminal, outputs the audio received from the other call terminal through the speaker, and transmits the video captured by the camera to the other call terminal, thereby enabling communication with the other call terminal. In the aforementioned call mode, the call terminal transmits the audio picked up by the microphone, the audio received from the other call terminal, and the video footage captured by the camera to the second recording device in parallel with the call using the aforementioned call protocol, and records them as a call record.

Citation Information

Patent Citations

  • Network communication exchange and network communication system using same

    JP2006114951A

  • Monitored video recording system, and monitored video reproducing and displaying method

    JP2009296207A

  • Intercom system

    JP2014229990A

  • Radio communication apparatus, radio communication system and data processing method

    JP2015119229A

  • Intercom device with television

    JP2020113957A