Communication System
The integration of intercoms, IP speakers, and surveillance cameras in a communication system addresses the lack of enhanced functionality, offering integrated voice calls, broadcasts, and event detection for improved security and communication.
Patent Information
- Application Number
- JP2021157990
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-12-10
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing communication systems lack integration and enhanced functionality by combining intercom and surveillance camera systems for improved security and communication capabilities.
A communication system that integrates intercoms, IP speakers, surveillance cameras, and video recording systems, allowing for voice calls, loudspeaker broadcasts, and event detection through audio and video analysis, with modes for call, loudspeaker, and monitoring standby operations, and utilizing AI for event detection.
Enhances communication systems with added value by enabling integrated voice calls, broadcasts, and event detection, providing improved security and operational efficiency.
Smart Images

Figure 0007783716000001 
Figure 0007783716000002 
Figure 0007783716000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a communication system that can integrate and operate in-house calls using intercoms and IP speakers, loudspeakers, video monitoring using a surveillance camera system, video recording, etc. [Background technology]
[0002] An intercom system is a system primarily used for internal calls. Typically, an intercom system includes one or more call terminals and one or more public address devices such as speakers. Each call terminal is equipped with a microphone and a speaker and is interconnected via a network. Each call terminal transmits input audio to another call terminal in accordance with a predetermined protocol such as the SIP protocol to perform internal calls. Each call terminal also transmits input audio to the public address device, amplifying the audio from the public address device.
[0003] Surveillance camera systems are systems primarily used for crime prevention. Typically, a surveillance camera system includes one or more surveillance cameras and a recorder to which each surveillance camera is connected. Each surveillance camera may be installed in a predetermined location, such as on a ceiling or wall, or may be mounted on a mobile object, such as a drone or automobile. The recorder receives footage captured by the surveillance camera and stores it on an internal or external recording medium, allowing live viewing by outputting the footage to an external display in real time.
[0004] An example of an intercom system is disclosed in Patent Document 1. An example of a surveillance camera system is disclosed in Patent Document 2. [Patent Document 1] JP 2006-114951 [Patent Document 2] Patent Publication No. 2009-296207 Summary of the Invention [Problem to be solved by the invention]
[0005] An object of the present invention is to provide a communication system with higher added value by, for example, combining an intercom system and a surveillance camera system. [Means for solving the problem]
[0006] The communication system includes a call terminal with a microphone and a loudspeaker connected to the call terminal via a network, and has selectable first and second loudspeaker modes. In the first loudspeaker mode, the call terminal generates an audio stream from the speaker's voice input from the microphone and transmits it to the loudspeaker, and the loudspeaker plays the received audio stream to perform loudspeaker broadcasting. In the second loudspeaker mode, the call terminal generates an audio file from the speaker's voice input from the microphone and downloads it to the loudspeaker, and the loudspeaker plays the downloaded audio file to perform loudspeaker broadcasting. [Effects of the Invention]
[0007] The present invention can provide a communication system with higher added value. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing a communication system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of the configuration of a communication system including a LAN and a WAN. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of a call terminal in a call system. [Figure 4] FIG. 4 is a diagram showing an example of a recording operation that differs depending on the operation mode. [Figure 5] FIG. 5 is a diagram showing an example of a recording operation that differs depending on the operation mode. [Figure 6]FIG. 6 is a flowchart showing an example of an event detection function between a communication system and a monitoring camera system. [Figure 7] FIG. 7 is a diagram showing an example of a database that associates the locations of call areas of a call system with the locations of monitoring areas of a monitoring camera system. [Figure 8] FIG. 8 is a diagram showing an example of a database that associates the locations of call areas of a call system with the locations of monitoring areas of a monitoring camera system. [Figure 9] FIG. 9 is a diagram showing an example of the configuration of a monitoring camera in a monitoring camera system. [Figure 10] FIG. 10 is a diagram showing an example of the configuration of a communication system according to an embodiment of the present invention, which includes an AI processing unit. [Figure 11] FIG. 11 is a diagram showing an example of the configuration of an IP speaker in a public address system. [Figure 12] FIG. 12 is a flowchart showing an example of the recorded broadcast function in the telephone system. [Figure 13] FIG. 13 is a flowchart showing an example of a download broadcast function between a call system and a loudspeaker system. [Figure 14] FIG. 14 is a flowchart showing an example of a broadcast function between the telephone system and the public address system. [Figure 15] FIG. 15 is a diagram showing an example of a configuration in which a paging gateway in a public address system is connected to an emergency signal source. [Figure 16] FIG. 16 is a flowchart showing an example of the operation flow of the paging gateway and the IP speaker in the coexistence function with the emergency broadcast. [Figure 17] FIG. 17 is a flowchart showing an example of an operation flow of a paging gateway in the coexistence function with an emergency broadcast. [Figure 18] FIG. 18 is a flowchart showing an example of an operation flow of an IP speaker in the coexistence function with an emergency broadcast. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 shows a communication system 1 according to the embodiment. The communication system 1 includes a call system 10, a system manager 20, a recorder 30, a surveillance camera system 40, and a loudspeaker system 50. The call system 10, the system manager 20, the recorder 30, the surveillance camera system 40, and the loudspeaker system 50 are connected to each other via a network.
[0010] The call system 10 includes one or more call terminals 11. The call terminal 11 is a terminal that performs voice calls and has the function of performing voice calls by transmitting and receiving voice data to and from other call terminals 11 via a network. Each call terminal 11 is assigned an identifier (e.g., an IP address) that uniquely identifies it on the network. During a voice call, the call terminal 11 transmits and receives audio data and video data in accordance with a protocol for IP calls. One example of a protocol for IP calls is SIP (Session Initiation Protocol), which is standardized by the IETF (Internet Engineering Task Force) and specified in RFC3261.
[0011] The system manager 20 is a control device that allows the call terminal 11 to perform transmission and reception processing of audio data and video data in accordance with an IP call protocol. The transmission and reception processing includes call control such as making calls. For example, the system manager 20 is a SIP server that performs call control in accordance with SIP. The system manager 20 is assigned an identifier (for example, an IP address) that uniquely identifies itself on the network.
[0012] The recorder 30 is a recording device equipped with a recording medium such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The recorder 30 conforms to the IP call protocol, just like the call terminal 11, and has a function of recording video data transmitted from the call terminal 11 in accordance with the IP call protocol. The recorder 30 is assigned an identifier (for example, an IP address) that uniquely identifies it on the network.
[0013] The surveillance camera system 40 is a system for monitoring a predetermined area (monitoring area) for security purposes, etc., and includes one or more surveillance cameras 41 and a recorder 42. Each surveillance camera 41 generates video data from a video signal acquired by an image sensor such as a CCD or CMOS, and transmits the video data externally. The surveillance camera 14 transmits streaming video in accordance with a transmission protocol for the surveillance camera system. One example of a transmission protocol for the surveillance camera system is a protocol conforming to the ONVIF (Open Network Video Interface Forum) standard. Each surveillance camera 41 is assigned an identifier (e.g., an IP address) that uniquely identifies it on the network. The surveillance camera 41 and the recorder 42 are recording devices equipped with a recording medium such as an HDD or SSD, and have the function of recording video data transmitted from the connected surveillance camera 41. The recorder 42 receives and records video data in accordance with the transmission protocol for the surveillance camera system, just like the surveillance camera 41. The recorder 42 is assigned unique identification information on the network (e.g., an IP address). The recorder 42 has a function for receiving audio input via a microphone, generating audio data from the audio input from the microphone, and transmitting the audio data to an external device. The recorder 42 transmits the audio data in accordance with a transmission protocol for a surveillance camera system. The surveillance camera system 40 may be managed in an integrated manner by a VMS (Video Management System). When managed in an integrated manner by a VMS, the surveillance camera 41 and the recorder 42 are controlled and managed by VMS software installed on a PC (Personal Computer). When managed in an integrated manner by a VMS, the recorder 42 may be a PC on which VMS software is installed and which is equipped with a recording medium such as an HDD or SSD, and in this case, the audio data of the audio input from the microphone equipped on the PC is transmitted to an external device.
[0014] The public address system 50 is a system installed in a predetermined area (public address area) for purposes such as evacuation guidance during a disaster and playing music such as background music, and includes one or more IP speakers 51. Each IP speaker 51 is assigned unique identification information (e.g., an IP address) on the network. The IP speaker 51 has the function of receiving, playing, and amplifying audio data in accordance with the same IP call protocol as the call terminal 11 and the same surveillance camera system transmission protocol as the recorder 42. The public address system 50 also includes a paging gateway 52 connected to the same network as the IP speaker 51. The paging gateway 52 is assigned unique identification information (e.g., an IP address) on the network. The paging gateway 52 has the function of receiving audio data via the network and transferring the audio data to the IP speaker 51 in accordance with the same IP call protocol as the call terminal 11 and the same surveillance camera system transmission protocol as the recorder 42. In other words, the paging gateway 52 can receive audio data intended for the IP speaker 51 and then transfer the audio data to the IP speaker 51, thereby broadcasting the audio data to multiple IP speakers. The paging gateway 52 may transfer the audio data to the IP speaker 51 by multicast or broadcast. The IP speaker 51 receives and plays back the digitized audio data via a wireless communication medium such as an Ethernet cable or cellular communication.
[0015] In this way, in the communication system 1, the call system 10, the system manager 20, the recorder 30, the surveillance camera system 40, and the loudspeaker system 50 are interconnected via a network, and it is possible to make voice calls between the call terminals 11 of the call system 10, to make voice broadcasts from the call terminals 11 of the call system 10 to the IP speakers 51 of the loudspeaker system 50, and to make voice broadcasts from the recorder 42 of the surveillance camera system 40 to the IP speakers 51 of the loudspeaker system.
[0016] As shown in FIG. 2, the communication system 1 may be configured via multiple local area networks (LANs) and a wide area network (WAN) such as the Internet that interconnects these networks. The call system 10, system manager 20, recorder 30, surveillance camera system 40, and loudspeaker system 50 are each configured on a LAN and interconnected. In this configuration, various types of communication are possible across the multiple LANs, such as voice calls between call terminals 11 in different call systems 10, voice broadcasts from a call terminal 11 in one call system 10 in one LAN to an IP speaker in a loudspeaker system 50 in another LAN, and voice broadcasts from a recorder 42 in a surveillance camera system 40 in one LAN to an IP speaker 51 in a loudspeaker system 50 in another LAN. When the call system 10 is configured on a LAN, the call terminal 11 is preferably connected to a router constituting the LAN via an Ethernet cable, and may transmit voice data via the Ethernet cable and be supplied with power over Ethernet (PoE) via the Ethernet cable. When the surveillance camera system 40 is configured on a LAN, the surveillance camera 41 and recorder 42 are preferably connected to a router that makes up the LAN with an Ethernet cable, and send and receive video data via the Ethernet cable, and are also supplied with PoE power via the Ethernet cable.When the public address system 50 is configured on a LAN, the IP speaker 51 and paging gateway 52 are preferably connected to a router that makes up the LAN with an Ethernet cable, and are also preferably connected to and receive audio data via the Ethernet cable, and are also supplied with PoE power via the Ethernet cable.
[0017] The following description will be given using an example in which SIP is used as the IP call protocol and the ONVIF protocol is used as the transmission protocol for the surveillance camera system.
[0018] As shown in Fig. 3, the communication terminal 11 includes a microphone 110, a speaker 111, a camera 112, a display unit 113, and an input unit 114. The microphone 110 is an electro-acoustic transducer whose sound collection direction faces outward from the housing of the communication terminal 11, and collects the voice of a speaker speaking to the communication terminal 11 and outputs an audio signal. The speaker 111 is an electro-acoustic transducer whose sound emission direction faces outward from the housing of the communication terminal 11, and outputs audio obtained by playing back audio data. The camera 112 is an imaging device including an imaging element such as a CCD image sensor or a CMOS image sensor, and its imaging direction faces outward from the housing of the communication terminal 11 in a direction that captures the figure of the speaker speaking to the communication terminal 11. The camera 112 generates and outputs video data of still images or videos from the video signal obtained by capturing an image. The display unit 113 is a display including a display element such as an LCD or an organic EL, and displays images captured by the camera 112 of the terminal 11 itself or another communication terminal 11, as well as other images related to GUIs and the like. The input unit 114 is an input device that accepts user operations, such as buttons or switches. The input unit 114 and the display unit 113 may be integrated to form a touch panel. The call terminal 11 can take various forms, such as a desk-top type, a wall-mounted type, or a wall-embedded type. The call terminal 11 is, so to speak, a call device called a speakerphone, an interphone, or an intercom.
[0019] (Operation mode) The communication terminal 11 operates in multiple operation modes. The operation modes include at least a standby mode, a call mode, and a loudspeaker mode. After the power of the communication terminal 11 is turned on, the communication terminal 11 operates in the standby mode while neither making calls to other communication terminals 11 nor transmitting loudspeakers to the IP speaker 51 or the paging gateway 52. The standby modes include a general standby mode and a monitoring standby mode. The general standby mode is a mode in which the microphone 110, the speaker 111, and the camera 112 are turned off and neither audio nor video is transmitted to the outside. The monitoring standby mode is a mode in which at least video is transmitted to the surveillance camera system 40 for security purposes, in which at least the camera 112 and optionally the microphone 110 are turned on and the speaker 111 is turned off. In the standby mode, the communication terminal 11 switches between the general standby mode and the monitoring standby mode manually in response to a user operation via the input unit 114 or automatically in accordance with predetermined conditions. For example, when in standby mode, the call terminal 11 operates in general standby mode by default, and switches to monitoring standby mode in response to a user operation from the input unit 114. For example, when in standby mode, the call terminal 11 operates in general standby mode by default, and switches to monitoring standby mode in response to the arrival of a predetermined time according to a predetermined schedule.
[0020] During standby mode, the call terminal 11 switches to a call mode or a loudspeaker mode either manually in response to a user operation via the input unit 114, or automatically in accordance with predetermined conditions. The call mode is a mode in which the call terminal 11 transmits and receives audio data to and from other call terminals 11, and a call is made between multiple call terminals 11. The loudspeaker mode is a mode in which the call terminal 11 transmits audio data to the IP speaker 51 and loudspeaks the audio from the IP speaker 51. For example, the call terminal 11 transitions to the call mode in response to an operation via the input unit 114 to select the destination of another call terminal 11 to be the call partner and make a call. For example, the call terminal 11 transitions to the call mode in response to receiving a call from another call terminal 11. For example, the call terminal 11 transitions to the call mode in response to a user operation via the input unit 114 to answer a call from another call terminal 11. For example, the communication terminal 11 transitions to loudspeaker mode in response to an operation via the input unit 114 to select the destination of the IP speaker 51 or the paging gateway 52 to be loudspeaked and make a call. In the call mode, the communication terminal 11 turns on the microphone 110, the speaker 111, and the camera 112 for a voice call (video call). In the loudspeaker mode, the communication terminal 11 turns on the microphone 110 for loudspeaker broadcasting. In the call mode, when communication with the call terminal 11 of the call destination is terminated and the call ends, the call terminal 11 ends the call mode and transitions to standby mode. In the loudspeaker mode, when communication with the IP speaker 51 or the paging gateway 52 is terminated, the call terminal 11 ends the loudspeaker mode and transitions to standby mode.
[0021] (Call function) In call mode, the call terminal 11 uses SIP to transmit audio data of the voice input from the turned-on microphone 110 and video data of the video captured by the camera 112 to another call terminal 11 (the call partner's call terminal 11) as the destination. Also, in call mode, the call terminal 11 plays back the audio data transmitted from the other call terminal 11 using SIP to output the audio from the speaker 111, and plays back the transmitted video data to display on the display unit 113. Also, in parallel with transmitting to the call terminal 11 of the call partner, the call terminal 11 transmits the audio data of the voice input from the microphone 110, the video data of the video captured by the camera 112, and the audio data and video data transmitted from the call terminal 11 of the call partner to the recorder 30 using SIP. The recorder 30 records the audio data and video data received according to SIP as a call record. When the call terminal 11 ends the call and ends the call mode, it responds by ending the transmission and reception of audio data and video data with the call terminal 11 of the other party and ending the transmission of the audio data and video data to the recorder 30. The recorder 30 starts recording when it starts receiving audio data and video data from the communication terminal 11 that has entered the call mode, ends recording when the call mode ends and the reception of audio data and video data from the communication terminal 11 stops, and saves the video data recorded during that time as a call record. The recorder 30 preferably generates one video file each time recording starts and ends, so that an individual video file is generated as a call record for each call made.
[0022] In the loudspeaker mode, the call terminal 11 transmits the audio data of the voice input from the microphone 110 that is turned on to the IP speaker 51 or the paging gateway 52 as the destination using SIP.
[0023] (Recording function) In monitoring standby mode, the communication terminal 11 generates video data of the video captured by the turned-on camera 112 and audio data of the audio input from the microphone 110, and constantly transmits them to the recorder 42 using the ONVIF protocol. The recorder 42 records the video data and audio data received according to the ONVIF protocol by recording them on a recording medium. The communication terminal 11 continues recording to the recorder 42 in monitoring standby mode even when the communication terminal 11 transitions to call mode or loudspeaker mode. The operation in monitoring standby mode will be described with reference to FIGS. 4 and 5. While set in monitoring standby mode, the communication terminal 11 constantly transmits audio data and video data to the recorder 42 to perform recording for security purposes. Even when the call mode is started during monitoring standby mode, the communication terminal 11 does not end the recording operation to the recorder 42. In the call mode, the communication terminal 11 continues to transmit audio data and video data to the recorder 42, while simultaneously transmitting and receiving audio data and video data to another call terminal 11 that is the call partner for the call, and transmitting the audio data and video data to the recorder 30 for call recording. When the call mode is ended, the call terminal 11 stops transmitting the audio data and video data to the recorder 30 for call recording, but continues transmitting the audio data and video data to the recorder 42. Similarly, even if the loudspeaker mode is started during the monitoring standby mode, the call terminal 11 does not end the recording operation to the recorder 42. In the loudspeaker mode, the communication terminal 11 continues to transmit the audio data and video data to the recorder 42, while simultaneously transmitting the audio data to the IP speaker 51 or the paging gateway 52 for loudspeaker broadcasting. When the loudspeaker mode is ended, the call terminal 11 stops transmitting the audio data to the IP speaker 51 or the paging gateway 52, but continues transmitting the audio data and video data to the recorder 42.
[0024] (Event detection function) The communication system 10 and the surveillance camera system 40 mutually detect events and respond to the detection. Here, an event is an occurrence that can be inferred from audio and video, and is an occurrence that can be inferred to have occurred in the location where the communication system 10 is installed (communication area) or the location where the surveillance camera system 40 is installed (monitoring area). One example of an event is a scream uttered by someone in the communication area or the monitoring area, which is an occurrence that can be inferred by analyzing audio, for example. Another example of an event is a suspicious person entering the communication area or the monitoring area, which is an occurrence that can be inferred by analyzing video, for example.
[0025] FIG. 6 is a flowchart showing the event detection function between the communication system 10 and the surveillance camera system 40. In the surveillance camera system 40, the surveillance camera 41 captures images of the surveillance area and detects events within the surveillance area from the captured images (S10). In S10, the surveillance camera 41 analyzes the video signal input from the image sensor to detect abnormalities within the surveillance area. For example, the surveillance camera system 40 has a motion detection function that compares multiple images captured by the surveillance camera 41 at predetermined time intervals (at a predetermined frame rate), identifies locations (pixel groups) where changes occur in the images, and detects moving objects such as people and vehicles at locations corresponding to these pixel groups. For example, the surveillance camera system 40 has an object detection function that stores image data of photographs of people, vehicles, etc. to be detected in advance, compares each of multiple images captured by the surveillance camera 41 at predetermined time intervals (at a predetermined frame rate) with the image data, and detects the presence of the person, vehicle, etc. within the surveillance area if a predetermined degree of match is found. The above detection function may be provided in the surveillance camera 41 or the recorder 42. In this way, when the surveillance camera system 40 detects an abnormality in the monitoring area from the video captured by the surveillance camera 41, it responds by transmitting a notification message to the call system 10 to notify the call system 10 of this (S11). In S11, the surveillance camera 41 or the recorder 42 transmits a call message to the call terminal 11 via the network. Here, the surveillance camera system 40 transmits the notification message to the call system 10 installed in the call area corresponding to the monitoring area in which the surveillance camera system 40 is installed. That is, in S11, the surveillance camera system 40 identifies a call system 10 having a call area in the same location as or closest to the monitoring area in which the abnormality was detected, and transmits a notification message to that call system 10.
[0026] The communication system 1 may include multiple communication systems 10, each with a different calling area, and surveillance camera systems 40, each with a different monitoring area. In this case, the communication system 1 maintains a database containing calling area location information indicating the location of the calling area of each communication system 10 and monitoring area location information indicating the location of the monitoring area of each monitoring camera system 40. FIG. 7 shows a database 70 as an example. The database 70 maintains the location of each of the multiple communication systems 10 and the location of each of the multiple monitoring camera systems 40. In the example shown in FIG. 5, in the database 70, the communication system 10-1 is associated with location AAAA, the communication system 10-2 is associated with location BBBB, the communication system 10-3 is associated with location CCCC, the monitoring camera system 40-1 is associated with location EEEE, the monitoring camera system 40-2 is associated with location AAAA, and the monitoring camera system 40-3 is associated with location BBBB. This indicates that each system is installed in the associated location. In the database 70, the information on each location may be manually entered by a user, for example, by an administrator of the communication system 10 or the surveillance camera system 40, or may be automatically entered by the system. As an example, the communication system 10 and the surveillance camera system 40 are equipped with a positioning system using a global navigation satellite system (GNSS) such as a global positioning system (GPS), and the database 70 is automatically generated with reference to the location information, such as longitude and latitude, acquired by the positioning system. As an example, the communication system 10 and the surveillance camera system 40 are equipped with an indoor positioning system using a beacon signal using Bluetooth or the like, RFID (Radio Frequency Identification), ultrasonic waves, geomagnetism, or UWB (Ultra Wide Band) signals, and the database 70 is automatically generated with reference to the indoor location information acquired by the indoor positioning system.As one example, user input of location information is accepted from input unit 114 provided in call terminal 11 of call system 10, and the accepted location information is input to database 70. As one example, user input of location information is accepted from input means (user interface such as buttons, switches, keyboards, etc.) provided in surveillance camera 41 or recorder 42 of surveillance camera system 40, and the accepted location information is input to database 70. Call system 10 and surveillance camera system 40 whose location information completely matches each other may be determined to be installed in the same location, and call system 10 and surveillance camera system 40 that are associated with substantially the same location information within a predetermined error range may be determined to be installed in the same location.
[0027] Fig. 8 shows another example of the database 70. In the example shown in Fig. 6, the database 70 holds a location associated with each communication terminal 11 in each communication system 10 and a location associated with each monitoring camera 41 in each monitoring camera system 40. In the example shown in Fig. 6, in the database 70, the communication terminal 11-1 in the communication system 10-1 is associated with location aaaa, the communication terminal 11-2 is associated with location bbbb, the communication terminal 11-3 is associated with location cccc, and the monitoring camera 41-1 in the monitoring camera system 40-1 is associated with location eeee, the monitoring camera 41-2 is associated with location aaaa, and the monitoring camera 41-3 is associated with location bbbb. This indicates that each communication terminal 11 and each monitoring camera 41 is installed in the associated location. In the database 70, information about each location may be manually entered by a user, for example, by an administrator of the communication terminal 11 or the surveillance camera 41, or may be automatically entered by the system. As an example, each communication terminal 11 and each surveillance camera 41 is equipped with a positioning system using GNSS such as GPS, and the location information, such as longitude and latitude, is acquired by the positioning system, and the database 70 is automatically generated by referring to the location information. As an example, each communication terminal 11 and each surveillance camera 41 is equipped with an indoor positioning system using beacon signals such as Bluetooth, RFID, ultrasound, geomagnetism, UWB signals, etc., and the indoor location information is acquired by the indoor positioning system, and the database 70 is automatically generated by referring to the indoor location information. As an example, user input of location information is accepted from an input unit 114 provided in the communication terminal 11, and the accepted location information is entered into the database 70. As an example, user input of location information is accepted from an input means (a user interface such as a button, switch, or keyboard) provided in the surveillance camera 41, and the accepted location information is entered into the database 70. A telephone terminal 11 and a surveillance camera 41 whose location information completely matches each other may be determined to be installed in the same location, and a telephone terminal 11 and a surveillance camera 41 that are associated with substantially the same location information within a predetermined error range may be determined to be installed in the same location.
[0028] In S11, the surveillance camera system 40 refers to the database 70 to identify a communication system 10 having a communication area that is the same as or closest to its own monitoring area, and transmits a notification message to the communication system 10. As an example, the surveillance camera system 40 identifies a communication system 10 that is associated with location information that substantially matches or is closest to its own location information, and transmits a notification message to the communication terminals 11 in the communication system 10 (for example, if there are multiple communication terminals 11, any, some, or all of the communication terminals 11). As an example, the surveillance camera system 40 identifies the surveillance camera 41 that detected the event in S10, identifies the communication terminal 11 that is associated with location information that substantially matches or is closest to the location information of the surveillance camera 41, and transmits a notification message to the identified communication terminal 11.
[0029] In the call system 10, the call terminal 11 (S12) that has received the notification message responds to the notification message and performs an operation to respond to the event (S13). An example of the event response operation is to start saving an audio signal input from the microphone 110 or a video signal captured by the camera 112. In S13, if the microphone 110 is OFF, the call terminal 11 turns the microphone 110 ON in response to the notification message and starts recording the audio input from the microphone 110. In S13, if the camera 112 is OFF, the call terminal 11 turns the camera 112 ON in response to the notification message and starts recording the video captured by the camera 112. The call terminal 11 may be provided with an internal recording medium such as an HDD or SSD, or may accept an external recording medium such as an SD card or USB memory (flash drive), and in S13, the recorded audio data or recorded video data may be saved on these recording media. In S13, the call terminal 11 may respond to the notification message by transitioning from the general standby mode to the monitoring standby mode, turning on the microphone 110 and the camera 112, and transmitting the recorded audio data and video data to the recorder 42, thereby starting storage in the recorder 42. In S13, the call terminal 11 may record audio and video in association with information related to the notification message to which it responded (for example, the identifier of the monitoring camera 41 or the recorder 42 that sent the notification message, the date and time the notification message was received, etc.). When the monitoring camera system 40 detects an event in a certain monitoring area, the call system 10 in the call area corresponding to the monitoring area can perform an operation in response to the event by steps S10 to S13. For example, when the monitoring camera 41 detects the intrusion of a suspicious person in a certain monitoring area, the call terminal 11 located in the same location as or nearby the monitoring area can start recording audio and video, and the situation of the suspicious person can be recorded.
[0030] Contrary to S10 to S14, the call system 10 monitors the call area using the microphone 110 and camera 112 provided in the call terminal 11. An event within the call area is detected from the audio input from the microphone 110 and the video captured by the camera 112 (S20), and the monitoring camera system 40 responds to this event (S23). In S20, the call system 10 analyzes the audio signal input from the camera 110 of the call terminal 11 to detect an abnormality within the call area. For example, the call system 10 has an abnormality detection function that performs frequency analysis of the audio signal input from the microphone 110 at predetermined time intervals and detects an abnormality when a specific frequency pattern corresponding to a specific sound such as a scream occurs, or monitors the level (amplitude) of the audio signal input from the microphone 110 at predetermined time intervals and detects an abnormality when a level above a predetermined level occurs. Furthermore, for example, the call system 10 has an abnormality detection function that analyzes the video signal input from the image sensor of the camera 112 to detect an abnormality within the call area. This is similar to the moving object detection function, object detection function, etc., of the monitoring camera system 40 described above. These detection functions are preferably provided in each communication terminal 11. In this way, when the communication system 10 detects an abnormality in the call area using the microphone 110 or camera 112 provided in the communication terminal 11, it responds by transmitting a notification message notifying the monitoring camera system 40 of the abnormality (S21). In S21, the communication terminal 11 that detected the abnormality transmits a call message to the monitoring camera 41 or recorder 42 via the network. Here, the communication system 10 transmits the notification message to the monitoring camera system 40 installed in the monitoring area corresponding to the call area in which the communication system 10 is installed. That is, in S21, similar to what the monitoring camera system 40 does in S11, the communication system 10 refers to the database 70 to identify a monitoring camera system 40 that has a monitoring area in the same location as or closest to the call area in which the abnormality was detected, and transmits a notification message to the monitoring camera system 40.As an example, the call system 10 identifies a surveillance camera system 40 associated with location information that substantially matches or is closest to its own location information, and transmits a notification message to the surveillance camera 41 (for example, any, some, or all of the surveillance cameras 41 if there are multiple surveillance cameras 41) or recorder 42 in the surveillance camera system 40. As an example, the call system 10 identifies the call terminal 11 that detected the event in S20, identifies the surveillance camera 41 associated with location information that substantially matches or is closest to the location information of the call terminal 11, and transmits a notification message to the identified surveillance camera 41.
[0031] As shown in FIG. 9, the surveillance camera 41 includes a camera 410 and a PTZ mechanism 411. The camera 410 is an imaging device including an imaging element such as a CCD image sensor or a CMOS image sensor, and is oriented so that the imaging direction is outside the housing of the surveillance camera 41 in a direction that captures part or all of the surveillance area in which the surveillance camera 41 is installed. The camera 410 generates and outputs still or moving image data from the video signal obtained by capturing the image. The PTZ mechanism 411 is connected to the camera 410, includes an actuator such as a motor, and is a mechanism that drives the camera 410 to pan, tilt, and zoom. The surveillance camera 41 is, for example, a camera that is attached to a wall or ceiling, and may be a box-type, bullet-type, dome-type, or other type of camera.
[0032] In the surveillance camera system 40, the surveillance camera 41 or recorder 42 (S22) that has received the notification message responds to the notification message and performs an operation to respond to the event (S23). One example of the event response operation is to drive the PTZ mechanism 411 of the surveillance camera 41 to pan, tilt, and zoom the camera 410, capture an image of the surveillance area, and store the video data in the recorder 42. In S23, if the surveillance camera 410 that has received the notification message is OFF, the surveillance camera 41 responds to the notification message by turning on the camera 410 and starting capture. More specifically, in S23, in response to the notification message, the surveillance camera 41 pans, tilts, and zooms in the direction of the communication terminal 11 that has sent the notification message, and starts capture. In S23, if the camera 410 of the connected surveillance camera 41 is OFF, the recorder 42 that has received the notification message responds to the notification message by turning on the camera 410 of the surveillance camera 41 and starting capture. More specifically, in S23, in response to the notification message, the recorder 42 causes the connected monitoring camera 41 to pan, tilt, and zoom in the direction of the communication terminal 11 that sent the notification message to capture an image. In S23, the monitoring camera 41 may store the video data in association with information about the response notification message (e.g., the identifier of the communication terminal 11 that sent the notification message, the date and time the notification message was received, etc.). When the communication system 10 detects an event in a certain communication area, the monitoring camera system 40 in the monitoring area corresponding to that communication area can perform an operation in response to the event. For example, when a scream is detected by the communication terminal 11 in a certain communication area, recording can be started by the monitoring camera 41 located in the same location as or nearby the communication area, and the situation in which the scream occurred can be recorded.
[0033] (Utilization of AI) The call system 10 may utilize AI (Artificial Intelligence) when detecting an event using the camera 110 or 112 provided in the call terminal 11. As shown in FIG. 10 , the call terminal 11 includes an AI processing unit 115. The call terminal 11 is also connected to a server 72 including an AI processing unit 720 via a WAN such as the Internet. The AI processing unit 115 is a processor that has a neural network and a trained learning model that has been generated in advance by machine learning (including deep learning). The server 72 is a web server accessible via the WAN and includes the AI processing unit 720. The AI processing unit 720 is a processor that has a neural network and a trained learning model that has been generated in advance by machine learning (including deep learning). The AI processing unit 720 is more expensive than the AI processing unit 115 but has advanced functionality and high processing power. The AI processing unit 115 is less expensive than the AI processing unit 720 but has lower functionality and processing power.
[0034] The AI processing unit 115 includes a learning model 115a. The learning model 115a is a model that receives input of audio data and video data and makes some kind of judgment based on the audio data and video data. The learning model 115a is configured to receive a large amount of samples of audio data and video data associated with a first event as training data, learn the features of the audio data and video data associated with the first event, and thereby make a judgment based on new input of audio data and video data. The AI processing unit 720 includes a learning model 720a. The learning model 720a is a model that receives input of audio data and video data and makes some kind of judgment based on the audio data and video data. The learning model 720a is configured to receive a large amount of samples of audio data and video data associated with a second event as training data, learn the features of the audio data and video data associated with the second event, and thereby make a judgment based on new input of audio data and video data. Here, the second event differs from the first event in that it is an event that requires higher computational power to be determined than the first event. That is, the process of determining the second event is a process with a higher load than the process of determining the first event. The determination process using the AI processing unit 115 is a local process that does not require communication via a WAN, and therefore is highly responsive to the call terminal 11, that is, the time required to obtain a determination result is short and the real-time nature is high. On the other hand, the determination process using the AI processing unit 720 can perform more advanced determinations, but requires communication via a WAN, and therefore is less responsive to the call terminal 11, that is, the time required to obtain a determination result is short and the real-time nature is low.
[0035] Taking into account the above-described characteristics of the AI processing unit 115 and the AI processing unit 720, the call terminal 11 appropriately uses the AI processing unit 115 and the AI processing unit 720. As an example, the call terminal 11 uses either the AI processing unit 115 or the AI processing unit 720 depending on the situation. As an example, the call terminal 11 uses the AI processing unit 115 and the AI processing unit 720 in cooperation with each other. When using the local AI processing unit 115, the call terminal 11 inputs audio data or video data to the AI processing unit 115 and receives a determination result by the learning model 115a. When using the AI processing unit 720 via a network, the call terminal 11 transmits audio data or video data to the AI processing unit 720 via a WAN and receives a determination result by the learning model 720a via the WAN.
[0036] For example, the first event is an audio-related event such as a scream, and the determination process for the first event is a process for determining an event such as a scream from audio data. The learning model 115a is configured to input audio data serving as a sample of an event such as a scream as training data, learn characteristics of the scream, and determine the occurrence of the scream from new input audio data. The second event is an image-related event such as the occurrence of a moving object, and the determination process for the second event is a process for determining an event such as the occurrence of a moving object from video data. The learning model 720a is configured to input video data serving as a sample of an event such as the occurrence of a moving object as training data, learn characteristics of the moving object, and determine the occurrence of the moving object from new input video data. The call terminal 11 may manually switch between determining an event such as a scream from audio data of input audio from the microphone 110 using the AI processing unit 115 and determining the occurrence of a moving object from video data of video captured by the camera 112 using the AI processing unit 720, depending on user input via the input unit 114. Furthermore, when the microphone 110 is turned on, the call terminal 11 may use the AI processing unit 115 to determine whether there is a scream or the like from the audio input to the microphone 110, and when the camera 112 is turned on, may use the AI processing unit 720 to determine whether there is a moving object or the like from the video captured by the camera 112. When both the microphone 110 and the camera 112 are turned on, the call terminal 11 may execute one or both of the determination process using the AI processing unit 115 and the determination process using the AI determination unit 720. Note that the first event being an audio-related event and the second event being an video-related event is merely an example and is not limited to this; conversely, the first event may be an video-related event and the second event may be an audio-related event.
[0037] For example, the determination of a first event is a first-stage determination related to audio or video, and the determination of a second event is a second-stage determination related to audio or video, which is a determination with a higher load than the first-stage determination. For example, the determination of a first event is a process of determining the occurrence of a moving object from video data, and the determination of a second event is a process of determining the occurrence of a moving object of a predetermined person or object, etc. That is, the determination of a first event may be a determination of the occurrence of some moving object regardless of what that object is, while the determination of a second event is a determination of a predetermined person (e.g., a wanted criminal) or object (e.g., a car with a specific license plate number) that is determined to be a specific person (e.g., a wanted criminal) or object (e.g., a car with a specific license plate number) that is determined to be a specific object to be determined. As an example, the communication terminal 11 uses the AI processing unit 115 and the AI processing unit 720 in stages. For example, the communication terminal 11 constantly drives the AI processing unit 115 and inputs the captured video of the camera 112 to the AI processing unit 115 at predetermined time intervals to perform the first event determination process. Next, when the AI processing unit 115 determines that a first event has occurred, the call terminal 11 responds by transmitting the video data on which the AI processing unit 115 has determined that the first event has occurred to the server 72, and then performs determination processing for a second event. By performing determination in stages in this manner, for example, the AI processing unit 115 can first determine that a moving object has occurred, and then the AI processing unit 720 can determine whether the moving object is a specific person or object. This allows for more efficient determination than performing determination processing by the AI processing unit 720 at all times.
[0038] (Public address function) In loudspeaker mode, the communication system 10 transmits audio data to the loudspeaker system 50, causing the IP speaker 51 to broadcast the audio. For example, each communication terminal 11 receives a user input from the input unit 114 specifying the destination IP speaker 51 or paging gateway 52, and transmits the audio data to one or more of the specified IP speakers 51 or paging gateways 52. When the communication system 10 wants to broadcast the same audio content from the IP speakers 51 in the loudspeaker system 50, it can transmit the audio data to the paging gateway 52. The paging gateway 52 then forwards the received audio data to all IP speakers 51 connected to the same network, causing all IP speakers 51 to play and output the audio data. As shown in FIG. 11 , the IP speaker 51 includes a speaker 510 and a memory unit 510. The speaker 510 is an electro-acoustic transducer whose sound output direction faces outward from the housing of the IP speaker 51, and outputs the audio obtained by playing back the audio data. The storage unit 511 is a storage medium that temporarily or non-temporarily stores the voice data received from the call system 10, and is a memory such as a ROM or RAM, or a storage such as an HDD or SSD, or a combination of these.
[0039] (Real-time broadcasting function) In loudspeaker mode, the communication system 10 has a real-time broadcasting function that instantly broadcasts audio input from the microphone 110 in real time. In real-time broadcasting, the communication system 10 transmits audio data of the audio input from the microphone 110 as an audio stream to a specified IP speaker 51 or paging gateway 52. For example, the communication terminal 11 generates audio data from the audio input from the microphone 110 and immediately transmits the audio stream to the IP speaker 51 or paging gateway 52 using a streaming protocol such as RTP (Realtime Transport Protocol). The IP speaker 51 immediately plays back and outputs the audio streams that it receives sequentially. The paging gateway 52 transfers the audio streams that it receives sequentially to the IP speaker 51 within the same network.
[0040] (Recorded broadcast function) The call system 10 has a recording broadcast function in the loudspeaker mode, which temporarily stores (records) input voice from the microphone 110 in the call terminal 11 and broadcasts the voice after the user (speaker) has listened to it again. FIG. 12 is a flowchart showing an example of the recording broadcast function. When the recording broadcast function is enabled, the call terminal 11 generates voice data from the input voice from the microphone 110, but temporarily stores the voice data in a recording medium instead of immediately transmitting it as an audio stream (S40). The call terminal 11 may include an internal recording medium such as an HDD or SSD, or may accept an external recording medium such as an SD card or USB memory (flash drive), and may store the recorded voice data in these recording media in S40. In S40, the call terminal 11 may automatically start recording when voice input from the microphone 110 begins as a result of the speaker starting to speak into the microphone 110. In S40, the call terminal 11 may accept a user operation from the input unit 114 to instruct the start of recording, and start recording in response to this. In S40, after starting recording, the call terminal 11 may automatically stop recording when the speaker stops speaking into the microphone 110 and the audio input from the microphone 110 ceases. In S40, after starting recording, the call terminal 11 may accept a user operation from the input unit 114 instructing to end recording, and end the recording in response to this. In S40, the call terminal 11 generates a series of audio data from the audio input from the start of recording to the end of recording, and stores the data in a storage medium. The call terminal 11 may retain the audio data stored in the storage medium in S40 for a certain period of time, and automatically delete the audio data from the recording medium after the certain period of time has elapsed. The call terminal 11 can play back the audio data stored in the storage medium for the certain period of time that the audio data is stored in the recording medium (S41). As an example, in S41, the call terminal 11 automatically plays back the audio data within the certain period of time in response to the end of recording in S40. As an example, in S41, the call terminal 11 accepts a user input instructing playback from the input unit 114 within a certain period of time, and in response to this, performs playback. In S41, the call terminal 11 outputs the played back audio from the speaker 111.After playback in S41, the call terminal 11 accepts, from the input unit 114, a user operation to instruct broadcasting or a user operation to instruct re-recording. In response to the user operation to instruct broadcasting, the call terminal 11 broadcasts the audio data played back in S41, that is, it transmits the audio stream of the audio data stored in the recording medium to the specified IP speaker 51 or paging gateway 52 (S42). Meanwhile, in response to a user operation to instruct re-recording, the call terminal 11 erases the audio data stored in the recording medium and returns to S40 to record the speaker's voice again. With the above-mentioned recording and playback functions, the user (speaker) can first play back and listen to what they have said to check it before broadcasting it, and if they are not satisfied with the content after listening, they can re-record it and broadcast it.
[0041] (Download broadcast function) In the loudspeaker mode, the communication system 10 has a download broadcasting function that downloads an audio file to the IP speaker 51 or the paging gateway 52 and plays back the downloaded audio file, rather than transmitting an audio stream using a streaming protocol. FIG. 13 is a flowchart showing an example of the download broadcasting function. When the download broadcasting function is enabled in the loudspeaker mode, the communication terminal 11 in the communication system 10 records audio input via the microphone 110, i.e., encodes the audio to generate audio data and stores it in a storage medium (S50). In S50, the communication terminal 11 may automatically start recording when audio input from the microphone 110 begins as a result of the speaker starting to speak into the microphone 110. In S50, the communication terminal 11 may accept a user operation from the input unit 114 instructing the start of recording, and start recording in response to this. In S50, the communication terminal 11 may automatically stop recording after starting recording when the speaker stops speaking into the microphone 110 and the audio input from the microphone 110 ceases. In S50, after starting recording, the call terminal 11 may accept a user operation from the input unit 114 to instruct to end recording, and may end the recording in response to this. In S50, the call terminal 11 generates a series of audio data from the audio input from the start of recording to the end of recording, and stores the data in a storage medium. In S50, the call terminal 11 may generate an audio file in accordance with a predetermined file format such as MP3, and store the audio file.
[0042] Next, the call terminal 11 downloads the voice data stored in the recording medium to the designated IP speaker 51 or paging gateway 52 (S51). In S51, the call terminal 11 may automatically start the download after completing the recording in S50. In S51, the call terminal 11 may accept a user operation from the input unit 114 instructing a download after completing the recording in S50, and start the download in response to this. In S51, the call terminal 11 transfers the voice data to the IP speaker 51 or paging gateway 52 using a file transfer protocol such as HTTP (Hypertext Transfer Protocol) or FTP (File Transfer Protocol).
[0043] The IP speaker 51 or the paging gateway 52 downloads the voice data and stores the downloaded voice data in a recording medium (S52). In S52, the IP speaker 51 stores the downloaded voice data in the storage unit 511. Like the IP speaker 51, the paging gateway 52 also includes a storage medium including a memory such as a ROM or RAM, or a storage such as an HDD or SSD, or a combination of these, and stores the voice data downloaded in the storage medium in S50. When the download of the voice data is completed, the IP speaker 51 or the paging gateway 52 transmits a message indicating this (download completion notification message) to the call terminal 11 (S53).
[0044] When audio data is stored in the IP speaker 51 or the paging gateway 52, the communication system 10 can instruct the loudspeaker system 50 to play back and broadcast the held audio data (S55 to S57). The communication terminal 11 accepts a user operation from the input unit 114 to instruct the built-in audio source broadcast, and in response thereto, transmits a request message (built-in audio source broadcast request message) instructing the broadcast of the built-in audio source to the designated IP speaker 51 or paging gateway 52 (S55). The IP speaker 51 receives the built-in audio source broadcast request message (S56), and in response thereto, plays back the audio data held in the storage unit 511 and outputs it from the speaker 510, thereby broadcasting it to the loudspeaker area (S57). If there are multiple pieces of audio data held in the storage unit 511, the communication terminal 11 accepts a user operation from the input unit 114 to select audio data, and as a result, information indicating the selected audio data (audio data identification information) is included in the built-in audio source broadcast request message. In S57, the IP speaker 51 refers to the audio data identification information included in the built-in audio source broadcast request message, and reads out from the storage unit 511 the audio data indicated by the audio data identification information, and plays back the audio data. When the paging gateway 52 receives the built-in audio source broadcast request message (S56), in response to this, it transfers the audio data stored in the storage medium to the IP speaker 51 within the same network, causing the IP speaker 51 to play back the audio data (S57). If there is a plurality of pieces of audio data stored in the storage medium, in S55 the call terminal 11 accepts a user operation to select audio data from the input unit 114, and information indicating the selected audio data (audio data identification information) is included in the built-in audio source broadcast request message. In S57, the paging gateway 52 refers to the audio data identification information included in the built-in audio source broadcast request message, and reads out from the storage medium the audio data indicated by the audio data identification information, and transfers the audio data. In S57, the paging gateway 52 may transfer the audio data by multicast or broadcast.With the download broadcasting function described above, audio data is first downloaded to the IP speaker 51 or paging gateway 52 and then played back, so even if the communication environment is not suitable for streaming playback, such as when the network bandwidth is narrow, good amplified broadcasting can be achieved without any sound dropouts.
[0045] As a variation of S50 to S57, when the download of the audio data is completed (S52), the loudspeaker system 50 may automatically play the audio data without waiting for an instruction from the communication system 10 (S57). In this variation, when the download of the audio data is completed (S52), the IP speaker 51 stores the audio data in a storage medium and transmits a download completion notification message, or automatically plays the audio data stored in the storage medium without transmitting a download completion notification message, thereby performing loudspeaker broadcasting (S57). Also, in this variation, when the download of the audio data is completed (S52), the paging gateway 52 stores the audio data in a storage medium and transmits a download completion notification message, or automatically transfers the audio data stored in the storage medium to the IP speaker 51, thereby causing the IP speaker 51 to play the audio data (S57).
[0046] (Broadcast function) When the communication system 10 is in loudspeaker mode and includes multiple IP speakers 51, the communication system 10 has a simultaneous broadcast function that broadcasts the same content from each IP speaker 51 at the same time. FIG. 14 is a flowchart showing an example of the simultaneous broadcast function. When the simultaneous broadcast function is enabled in loudspeaker mode, the communication terminal 11 in the communication system 10 records audio input via the microphone 110, i.e., encodes the audio to generate audio data, and stores the audio data in a storage medium (S60). The recording operation in S60 is the same as the recording operation in S50. Next, the communication terminal 11 downloads the audio data stored in the storage medium to the multiple IP speakers 51 in the loudspeaker system 50 (S61). In S61, the communication terminal 11 accepts a user operation from the input unit 114 to select two or more IP speakers 51, and in response, transmits the audio data to the selected IP speakers 51. In S61, the call terminal 11 may receive a user operation from the input unit 114 instructing simultaneous broadcasting, and in response thereto, transmit audio data to all IP speakers 51 in the loudspeaker system 50. In S61, the call terminal 11 may automatically start downloading after completing recording in S60. In S61, the call terminal 11 may receive a user operation from the input unit 114 instructing downloading after completing recording in S60, and in response thereto, start downloading. In S61, the call terminal 11 downloads the audio data by transferring it to the IP speaker 51 using a file transfer protocol such as HTTP or FTP.
[0047] Each IP speaker 51 downloads the audio data and stores the downloaded audio data in a recording medium (S62). In S62, the IP speaker 51 stores the downloaded audio data in the storage unit 511. When the IP speaker 51 completes the download of the audio data, it transmits a message indicating this (download completion notification message) to the call terminal 11 (S63). After starting the transmission of the audio data (S61), the call system 10 waits for the reception of download completion notification messages from all of the destination IP speakers 51 (S64, S65). When the call terminal 11 receives the download completion notification messages from all of the IP speakers 51 (S65: Yes), the call system 10 automatically instructs the loudspeaker system 50 to play and broadcast the downloaded audio data (S66). In S66, in response to receiving the download completion notification messages from all of the IP speakers 51 that downloaded the audio data in S61 (S65: Yes), the call terminal 11 automatically transmits a broadcast instruction message to instruct all of these IP speakers 51 to broadcast.
[0048] Each IP speaker 51 receives the broadcast instruction message (S67) and, in response, plays the audio data downloaded and stored in storage unit 511 and outputs it from speaker 510, thereby broadcasting to the loudspeaker area (S68). Steps S60 to S68 ensure broadcasting capability when multiple IP speakers 51 simultaneously broadcast within the loudspeaker area. In other words, if multiple IP speakers 51 individually receive, play, and broadcast an audio data stream, differences in the speed at which the audio stream is received among the IP speakers 51 may occur depending on the network conditions between the call system 10 and loudspeaker system 50 (the network conditions between the call terminal 11 and each IP speaker 51), which could result in a time lag in the sound output among the IP speakers 51. Furthermore, even in the case of a method in which voice data is converted into a file in the call system 10 and downloaded to the loudspeaker system 50 before being played back, there is a possibility that, depending on the network conditions between the call system 10 and the loudspeaker system 50 (the network conditions between the call terminal 11 and each IP speaker 51), there may be differences in the time it takes for the download to be completed among the IP speakers 51, resulting in a time lag in the sound output among the IP speakers 51. In contrast, according to S60 to S68, once the download of the voice data has been completed on all IP speakers 51 that are the broadcast targets, instructions to broadcast the downloaded voice data are issued all at once, so that all IP speakers 51 can start broadcasting at approximately the same time.
[0049] (Coexistence with emergency broadcast function) Public address systems 50 can be installed in commercial facilities such as office buildings, factories, and shopping malls. Depending on their size, these facilities may be required by fire safety regulations to install emergency broadcast equipment. Emergency broadcast equipment amplifies emergency broadcast audio and outputs it from multiple emergency broadcast speakers installed on each floor of the facility. When public address systems 50 are installed in facilities equipped with emergency broadcast equipment, IP speakers 51 are also installed on each floor. Care must be taken to ensure that the IP speakers 51 do not compete with the emergency broadcast audio output from the emergency broadcast speakers when the emergency broadcast equipment is activated. In other words, according to fire safety regulations, when the emergency broadcast equipment is activated, the emergency broadcast audio output from the emergency broadcast speakers must be clearly distinguishable from other sounds. Therefore, when the emergency broadcast equipment is activated, a mechanism is required to prevent audio from being output from other speakers installed in the same location as the emergency broadcast speaker. Conventionally, so-called commercial broadcast equipment installed in the same location as the emergency broadcast equipment uses a power cut-off relay that cuts off the power supply upon receiving a signal from the emergency broadcast equipment or an automatic fire alarm system. However, in the public address system 50, the IP speaker 51 may be connected to a router via an Ethernet cable and receive PoE power. In such a configuration, the router must be turned off to cut off power to the IP speaker 51, but devices other than the IP speaker 51 may also be connected to the router, and turning off power to the router itself may not be desirable. In light of this situation, the public address system 50 has a broadcast control function that allows it to coexist with emergency broadcast equipment without cutting off the power.
[0050] As shown in Fig. 15, in the public address system 50, the paging gateway 52 is connected to an emergency signal source 80 independently of the path connecting it to the communication system 10 and the like via the network. The emergency signal source 80 is a signal source that issues a signal indicating the occurrence of an emergency such as a fire, which activates emergency broadcast equipment installed in the facility where the public address system 50 is installed. The emergency signal source 80 is, for example, an automatic fire alarm system or an emergency broadcast equipment. The emergency signal source 80 is configured to issue an emergency signal to the outside when it detects the occurrence of an emergency such as a fire, and the paging gateway 52 receives this emergency signal.
[0051] 16 is a flowchart showing an example of a broadcast control function for coexistence with emergency broadcast equipment. The paging gateway 52 constantly monitors whether an emergency signal has been received from the emergency signal source 80 at predetermined time intervals (S80), and if an emergency signal has been received (S80: Yes), in response, transmits a message indicating that broadcasting is prohibited (broadcast prohibition message) to all IP speakers 51 within the same network (S81). The broadcast prohibition message may be transmitted to the IP speakers 51 using multicast or broadcast.
[0052] Each IP speaker 51 normally operates with its operation mode set to broadcast-enabled mode by default (S82). In the broadcast-enabled mode, the IP speaker 51 is capable of timely broadcasting, and plays and outputs audio data according to, for example, the real-time broadcasting function, recorded broadcasting function, download broadcasting function, and simultaneous broadcasting function described above. For example, in the broadcast-enabled mode, the IP speaker 51 plays and outputs audio streams that are sequentially received according to the real-time broadcasting function or recorded broadcasting function. For example, in the broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to a broadcast instruction received according to the download broadcasting function (S56-S57). For example, in the broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to a broadcast instruction received according to the simultaneous broadcasting function (S67-S68). On the other hand, when the IP speaker 51 receives a broadcast prohibition message from the paging gateway 52 (S83), it sets its operation mode to broadcast prohibition mode in response to this (S84). In the broadcast prohibition mode, the IP speaker 51 cannot perform loudspeaker broadcasting and does not play or output audio data. For example, in the broadcast prohibition mode, even if the IP speaker 51 receives an audio stream according to the real-time broadcasting function or the recorded broadcasting function, the IP speaker 51 does not play the audio stream. For example, in the broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction according to the download broadcasting function (S56), the IP speaker 51 does not play audio data. For example, in the broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction according to the simultaneous broadcasting function (S67), the IP speaker 51 does not play audio data. By controlling in the above manner, when it becomes necessary to make an emergency broadcast, the operating mode of the IP speaker 51 can be transitioned to the broadcast prohibition mode, thereby preventing interference with the emergency broadcast.
[0053] 17 and 18 are flowcharts showing an example of a broadcast control function for coexistence with emergency broadcast equipment. The paging gateway 52 constantly monitors whether an emergency signal has been received from the emergency signal source 80 at predetermined time intervals (S91), and unless an emergency signal is received, constantly transmits a message indicating broadcast permission (broadcast permission message) to all IP speakers 51 within the same network (S91: No, S90). The broadcast permission message may be continuously transmitted to the IP speakers 51 at predetermined time intervals using multicast or broadcast. When the paging gateway 52 detects the reception of an emergency signal from the emergency signal source 80 (S91: Yes), it immediately and automatically stops transmitting the emergency signal in response (S92).
[0054] Each IP speaker 51 constantly monitors at predetermined intervals whether a broadcast permission message has been received from the paging gateway 52 (S95), and operates in broadcast-enabled mode as long as a broadcast permission message has been received (S96). In broadcast-enabled mode, the IP speaker 51 is capable of timely broadcasting, and plays and outputs audio data according to, for example, the real-time broadcasting function, recorded broadcasting function, download broadcasting function, and simultaneous broadcasting function described above. For example, in broadcast-enabled mode, the IP speaker 51 plays and outputs audio streams that are sequentially received according to the real-time broadcasting function or recorded broadcasting function. For example, in broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to a broadcast instruction received according to the download broadcasting function (S56-S57). For example, in broadcast-enabled mode, the IP speaker 51 plays and outputs audio data in response to a broadcast instruction received according to the simultaneous broadcasting function (S67-S68). On the other hand, when each IP speaker 51 no longer receives a broadcast permission message from the paging gateway 52 (S95: No), it immediately and automatically sets its operating mode to broadcast prohibition mode in response to this (S97). In broadcast prohibition mode, the IP speaker 51 cannot perform loudspeaker broadcasting and does not play or output audio data. For example, in broadcast prohibition mode, even if the IP speaker 51 receives an audio stream in accordance with the real-time broadcast function or the recorded broadcast function, the IP speaker 51 does not play the audio stream. For example, in broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction in accordance with the download broadcast function (S56), the IP speaker 51 does not play audio data. For example, in broadcast prohibition mode, even if the IP speaker 51 receives a broadcast instruction in accordance with the simultaneous broadcast function (S67), the IP speaker 51 does not play audio data.
[0055] As a variation of S95 to S97, the IP speaker 51 may confirm receipt of the broadcast permission message (S95) not constantly but only when broadcasting. That is, for example, in response to starting reception of an audio stream according to the real-time broadcasting function or the recorded broadcasting function, the IP speaker 51 may confirm receipt of the broadcast permission message (S95) before starting playback of the audio stream. In this case, the IP speaker 51 may buffer the audio stream and start playback of the buffered audio stream only when it is confirmed that the broadcast permission message has been received (S95: Yes). Also, for example, in response to receiving a broadcast instruction according to the download broadcasting function, the IP speaker 51 may confirm receipt of the broadcast permission message (S95) before starting playback of audio data according to the broadcast instruction. In this case, the IP speaker 51 may start playback of the audio data (S57) only when it is confirmed that the broadcast permission message has been received (S95: Yes). Furthermore, for example, in response to receiving a broadcast instruction according to the simultaneous broadcast function, the IP speaker 51 confirms receipt of a broadcast permission message (S95) before starting to play back audio data according to this broadcast instruction (S68). In this case, the IP speaker 51 may start playing back audio data (S68) only if it is confirmed that the broadcast permission message has been received (S95: Yes).
[0056] As described above, when the paging gateway 52 does not receive an emergency signal (when the emergency broadcast is not in operation), each IP speaker 51 plays and broadcasts audio data in response to a constant broadcast permission message sent from the paging gateway 52, and when reception of the broadcast permission message ceases, it determines that the audio data will interfere with the emergency broadcast and does not play or broadcast the audio data. With the broadcast control shown in S80 to S84, if the broadcast prohibition message sent from the paging gateway 52 in S81 does not reach the IP speaker 51 due to packet loss or the like, there is a risk that the operating mode of the IP speaker 51 will not be able to transition to the broadcast prohibition mode. The broadcast control shown in S90 to S97 can more reliably prevent interference with the emergency broadcast.
[0057] (supplement) The call terminal 11 includes a processing unit including a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), and a memory that stores computer programs that can be executed by the processing unit. The processing unit executes the computer programs stored in the memory, thereby realizing various functions of the call terminal 11. The computer programs stored in the memory include program instructions that execute the various functions described above, such as a call function, a recording function, an event detection function, an AI-based determination function, a loudspeaker function, a real-time broadcast function, a recorded broadcast function, a download broadcast function, and a simultaneous broadcast function.
[0058] The recorder 30 includes a processing unit including a CPU, an MPU, etc., and a memory that stores computer programs executable by the processing unit. The processing unit executes the computer programs stored in the memory to realize various functions of the recorder 30. The computer programs stored in the memory include program instructions that execute the various functions described above, such as the recording function.
[0059] The surveillance camera 41 includes a processing unit including a CPU, an MPU, etc., and a memory that stores computer programs executable by the processing unit. The processing unit executes the computer programs stored in the memory to realize various functions of the surveillance camera 41. The computer programs stored in the memory include program instructions that execute the various functions described above, such as the event detection function.
[0060] The recorder 42 includes a processing unit including a CPU, an MPU, etc., and a memory that stores computer programs that can be executed by the processing unit. The processing unit executes the computer programs stored in the memory to realize various functions of the recorder 42. The computer programs stored in the memory include program instructions that execute the various functions described above, such as the recording function and the event detection function.
[0061] The IP speaker 51 includes a processing unit including a CPU, an MPU, etc., and a memory that stores computer programs that can be executed by the processing unit. The various functions of the IP speaker 51 are realized by the processing unit executing the computer programs stored in the memory. The computer programs stored in the memory include program instructions that execute the various functions described above, such as the loudspeaker function, real-time broadcast function, recorded broadcast function, download broadcast function, simultaneous broadcast function, and function for coexistence with emergency broadcasts.
[0062] The paging gateway 52 comprises a processing unit including a CPU, an MPU, etc., and a memory that stores computer programs that can be executed by the processing unit. The various functions of the paging gateway 52 are realized by the processing unit executing the computer programs stored in the memory. The computer programs stored in the memory include program instructions that execute the various functions described above, such as the loudspeaker function, real-time broadcast function, recorded broadcast function, download broadcast function, simultaneous broadcast function, and function for coexistence with emergency broadcast. [Explanation of symbols]
[0063] 1. Communication System 10. Call System 20 System Manager 30 Recorder 40 Surveillance Camera System 50 Public address system
Claims
1. a call terminal having a speaker and a microphone; one or more loudspeakers connected to the call terminal via a network; having a selectable first loudspeaker mode and a selectable second loudspeaker mode; the call terminal has a recording function for recording the voice of the speaker input from the microphone, In the first loudspeaker mode, the communication terminal transmits the audio data recorded by the recording function and held to the loudspeaker device via a network using a streaming protocol, and performs loudspeaker broadcasting by causing the loudspeaker device to play back the received audio data; an intercom system in which, in the second loudspeaker mode, the call terminal generates an audio file from audio data recorded by the recording function and stored, downloads the audio file to the loudspeaker via a network using a file transfer protocol, and plays the downloaded audio file on the loudspeaker to perform loudspeaker broadcasting, In the first loudspeaker mode, after recording using the recording function, the call terminal makes the audio data stored by the recording possible to be played back from the speaker within a predetermined time before transmitting the audio data to the loudspeaker, and after playing back from the speaker within the predetermined time, transmits the audio data to the loudspeaker if an input operation instructing loudspeaker broadcasting is received from the speaker.
2. The intercom system according to claim 1, If there are two or more public address systems, In the second loudspeaker mode, the call terminal downloads the audio file to the two or more loudspeaker devices, and upon receiving a download completion notification from all of the two or more loudspeaker devices, causes the audio file to be played on all of the two or more loudspeaker devices.
3. Equipped with a speaker and a microphone, it is connected to one or more public address systems via a network. A call terminal, having a selectable first loudspeaker mode and a selectable second loudspeaker mode; a recording function for recording the voice of the speaker input from the microphone; In the first loudspeaker mode, the audio data recorded by the recording function is transmitted to the loudspeaker device via a network using a streaming protocol, thereby causing the loudspeaker device to play back the received audio data; In the second loudspeaker mode, an audio file is generated from the audio data recorded by the recording function and stored, and the audio file is downloaded to the loudspeaker via the network using a file transfer protocol, thereby causing the loudspeaker to play the downloaded audio file; In the first loudspeaker mode, after recording using the recording function, the voice data held by the recording can be played back from the speaker within a predetermined time before transmitting the voice data to the loudspeaker, and after playing back from the speaker within the predetermined time, the calling terminal transmits the voice data to the loudspeaker when an input operation instructing loudspeaker broadcasting is received from the speaker.
4. The call terminal according to claim 3, If there are two or more public address systems, In the second loudspeaker mode, the call terminal downloads the audio file to the two or more loudspeaker devices, and upon receiving a download completion notification from all of the two or more loudspeaker devices, plays the downloaded audio file on all of the two or more loudspeaker devices.
5. The intercom system according to claim 1, An intercom system in which the call terminal, in the first loudspeaker mode, erases the audio data recorded by the recording function and stored after the predetermined time has elapsed.
6. The call terminal according to claim 3, The communication terminal erases the voice data recorded and held by the recording function after the predetermined time has elapsed in the first loudspeaker mode.
Citation Information
Patent Citations
Network communication exchange and network communication system using same
JP2006114951A
IP broadcast transmitter, IP broadcast receiver, IP broadcast system equipped with them, and IP broadcast method
JP2006166238A
Audio data transmission system
JP2009267913A
Monitored video recording system, and monitored video reproducing and displaying method
JP2009296207A