Live audio transmission device and live audio transmission method

The live audio transmission device automatically converts scene images into audio commentary for timely transmission to headquarters, addressing the challenge of verbal communication limitations by using different transmission times based on image features, ensuring efficient and clear situational updates.

WO2026062948A1PCT designated stage Publication Date: 2026-03-26JVC KENWOOD CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Team members such as police officers or firefighters often cannot verbally convey the situation at the scene due to their circumstances, necessitating a live audio transmission device and method that can automatically transmit the scene's situation by voice.

Method used

A live audio transmission device and method that includes a situation explanation audio acquisition unit to capture images and convert them into audio commentary, wirelessly transmitting this commentary to a predetermined device using wireless communication, with different transmission times based on the presence or absence of matching features in the captured images.

Benefits of technology

Enables automatic and efficient transmission of detailed or concise audio commentary to headquarters based on the presence or absence of specific features in the captured images, reducing the need for manual communication and ensuring clear situational updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025014621_26032026_PF_FP_ABST
    Figure JP2025014621_26032026_PF_FP_ABST
Patent Text Reader

Abstract

In the present invention, a situation explanation audio acquisition unit (smartphone 3), if a feature which matches preset feature information is included in a captured image obtained by photographing the area forward of a user located at a scene, acquires first situation explanation audio for explaining an on-scene situation in a first time on the basis of the captured image and, if a feature which matches the feature information is not included in the captured image, acquires second situation explanation audio for explaining the on-scene situation in a second time shorter than the first time on the basis of the captured image. A wireless transmission unit (wireless transmission / reception unit 11) wirelessly transmits the first or second situation explanation audio acquired by the situation explanation audio acquisition unit to a prescribed device (main unit 60).
Need to check novelty before this filing date? Find Prior Art

Description

Live audio transmission device and live audio transmission method

[0001] The present disclosure relates to a live audio transmission device and a live audio transmission method.

[0002] Patent Document 1 describes that a police officer wears a wearable camera and records an image taken at the time of an incident.

[0003] Japanese Unexamined Patent Application Publication No. 2017-60029

[0004] There are cases where a team member such as a police officer or a firefighter wants to verbally convey the situation at the scene to a predetermined device (for example, the headquarters). However, since team members are often in situations where they cannot speak, the emergence of a live audio transmission device and a live audio transmission method that can automatically transmit the situation at the scene by voice is desired.

[0005] A first aspect of one or more embodiments provides a live audio transmission device including a situation explanation audio acquisition unit that acquires a first situation explanation audio for explaining the situation at the scene at a first time based on the captured image when the captured image of the front of the user located at the scene includes a feature that matches preset feature information, and acquires a second situation explanation audio for explaining the situation at the scene at a second time shorter than the first time based on the captured image when the captured image does not include a feature that matches the feature information, and a wireless transmission unit that wirelessly transmits the first or second situation explanation audio acquired by the situation explanation audio acquisition unit to a predetermined device by a first wireless communication method.

[0006] A second aspect of one or more embodiments provides a live commentary audio transmission method comprising: a camera attached to a user located at a site photographs the area in front of the user, and when the image captured by the camera contains features that match pre-set feature information, a situation commentary audio acquisition unit acquires a first situation commentary audio that describes the situation at the site in a first time based on the captured image; when the image does not contain features that match the feature information, the situation commentary audio acquisition unit acquires a second situation commentary audio that describes the situation at the site in a second time shorter than the first time based on the captured image; and a wireless transmission unit wirelessly transmits the first or second situation commentary audio to a predetermined device.

[0007] According to one or more embodiments of the live commentary audio transmission device and live commentary audio transmission method, the situation at the site can be automatically transmitted by voice.

[0008] Figure 1 is a block diagram of a live commentary audio transmission device according to the first embodiment. Figure 2 is a diagram showing an example of a state in which a police officer, who is a user, is wearing the live commentary audio transmission device according to the first or second embodiment. Figure 3 is a block diagram of a first configuration example of the live commentary audio transmission device according to the first or second embodiment. Figure 4 is a block diagram of a second configuration example of the live commentary audio transmission device according to the first or second embodiment. Figure 5 is a block diagram of a third configuration example of the live commentary audio transmission device according to the first or second embodiment. Figure 6 is a flowchart showing the processing performed by the live commentary audio transmission device according to the first embodiment. Figure 7 is a block diagram of a live commentary audio transmission device according to the second embodiment. Figure 8A is a diagram showing a state in which the timing is controlled to avoid collision when both of the two live commentary audio transmission devices according to the second embodiment transmit situational commentary audio for the first time to headquarters. Figure 8B is a diagram showing a state in which the timing is controlled to avoid collision when one of the two live commentary audio transmission devices according to the second embodiment transmits situational commentary audio for the first time to headquarters and the other transmits situational commentary audio for the second time to headquarters. Figure 9 is a flowchart showing the processing performed by the live commentary audio transmission device according to the second embodiment.

[0009] The following describes the live commentary audio transmission device and live commentary audio transmission method according to the first and second embodiments with reference to the attached drawings. In the first and second embodiments, the user is a police officer. The user is not limited to a police officer; it may be a firefighter or a member of the military.

[0010] <First Embodiment> Figure 1 shows a live commentary audio transmission device 100 according to the first embodiment. In Figure 1, the image analysis server 50 and the headquarters 60 are external components of the live commentary audio transmission device 100. The headquarters 60 is an example of a predetermined device that can communicate with the live commentary audio transmission device 100. Before describing the configuration and operation of the live commentary audio transmission device 100, Figures 2 to 5 will be used to explain how the user wears the live commentary audio transmission device 100 on their body.

[0011] As shown in Figure 2, user 40 is wearing the radio 1 on his belt. The radio 1 is connected to an auxiliary unit 2 by a wire, and the auxiliary unit 2 is worn on the upper left side of his chest. User 40 is wearing a smartphone 3 near his right hip. The radio 1 and smartphone 3 are connected by a wire. The auxiliary unit 2 and smartphone 3 are also connected by a wire. The lens 21L of the camera 21, which will be described later, is attached to a predetermined position on the auxiliary unit 2. The mounting positions of the radio 1, auxiliary unit 2, and smartphone 3 on user 40 are examples and are not limited to the mounting positions shown in Figure 2.

[0012] Figure 3 shows the wireless device 1, auxiliary device 2, and smartphone 3 shown in Figure 2. Figure 3 is considered the first configuration example. Wireless device 1 communicates wirelessly with headquarters 60 using a half-duplex communication method. Auxiliary device 2 has a camera 21, a speaker 22, and a microphone 23. Auxiliary device 2 can be configured as a speaker-microphone, which has a speaker and a microphone, with a built-in camera, and a central processing unit (CPU) controlling each part. Smartphone 3 communicates wirelessly with image analysis server 50 using a mobile phone network. Image analysis server 50 is an example of an image analysis device.

[0013] The image analysis server 50 is equipped with artificial intelligence known as multimodal AI. Multimodal AI is an AI that can handle multiple different types of information together, for example, it can combine and associate different types of information such as images, audio, and text for processing.

[0014] Instead of the first configuration example shown in Figure 3, the second configuration example shown in Figure 4 may be adopted. In the second configuration example, an auxiliary device 2' having a speaker 22 and a microphone 23 is used instead of the auxiliary device 2. The auxiliary device 2' may be a speaker microphone. The smartphone 3 has a camera 31. In the second configuration example, the camera 31 built into the smartphone 3 is used instead of the camera 21. When using the second configuration example, it is preferable for the user 40 to wear the smartphone 3 on their upper body, such as on their right chest.

[0015] The third configuration shown in Figure 5 may be used instead of the first configuration example shown in Figure 3 and the second configuration example shown in Figure 4. In Figure 5, the smartphone 3 has a camera 31, a speaker 32, a microphone 33, a central processing unit (hereinafter referred to as CPU) 34, and a storage unit 35 that stores a wireless application program (wireless AP). By having the CPU 34 execute the wireless AP, the smartphone 3 can function as a wireless device that communicates wirelessly with the headquarters 60 using a half-duplex communication method via a mobile phone network.

[0016] The live commentary audio transmission device 100 shown in Figure 1 is an example where the configuration shown in Figure 3 is adopted, but any of the configurations shown in Figures 3 to 5 may be used. Furthermore, the live commentary audio transmission device 100 may employ a configuration different from that shown in Figures 3 to 5.

[0017] In Figure 1, the smartphone 3 has a feature information setting unit 301, a feature extraction unit 302, a situation explanation instruction unit 303, a situation explanation text receiving unit 304, and a voice conversion unit 305. The feature information setting unit 301, the feature extraction unit 302, the situation explanation instruction unit 303, the situation explanation text receiving unit 304, and the voice conversion unit 305 are composed of, for example, a CPU. As an example, suppose user 40 is a police officer and is tracking a suspect. The suspect's characteristics are that he is a man wearing black pants, a black jacket, a hat, sunglasses, and a mask.

[0018] The radio 1 has a radio transmitting / receiving unit 11. The radio transmitting / receiving unit 11 includes a radio transmitting unit and a radio receiving unit. The radio receiving unit receives characteristic information of the suspect transmitted from headquarters 60. The characteristic information setting unit 301 sets the characteristic information by saving the received characteristic information of the suspect. The characteristic information may be a photograph of the suspect, or it may be information describing the suspect's characteristics. Information describing the suspect's characteristics may be information indicating that the suspect is male, wearing black pants, a black jacket, a hat, sunglasses, and a mask.

[0019] Camera 21 generates an image of the area in front of the user 40 located at the site and supplies it to the feature extraction unit 302. Smartphone 3 transmits the captured image to the image analysis server 50 via the mobile phone network. If the captured image contains features that match the feature information set in the feature information setting unit 301, the feature extraction unit 302 extracts the features that match the feature information from the captured image and supplies them to the situation explanation instruction unit 303. The feature extraction unit 302 can extract features that match the feature information by comparing the feature information with the captured image using known pattern recognition.

[0020] The situation explanation instruction unit 303 transmits a first instruction signal to the image analysis server 50 via the mobile phone network, instructing it to generate text data that explains the situation at the site based on the extracted features and the transmitted captured images, within a first time frame. In addition to the captured images, the microphone 23 may also transmit audio to the image analysis server 50, and the image analysis server 50 may be configured to analyze the situation at the site based on the captured images and audio.

[0021] If the feature extraction unit 302 does not contain any features that match the feature information set in the feature information setting unit 301, it notifies the situation explanation instruction unit 303 accordingly. For example, the feature extraction unit 302 extracts feature quantities from the captured image such as gender, trouser color, jacket color, presence or absence of a hat, presence or absence of sunglasses, and presence or absence of a mask, and notifies the situation explanation instruction unit 303 whether or not the feature information is included. The situation explanation instruction unit 303 transmits a second instruction signal via the mobile phone network to the image analysis server 50, instructing it to analyze the situation at the scene based on the transmitted captured image and generate text data that explains the situation in a second time, which is shorter than the first time.

[0022] The first and second times do not need to be fixed; the first time may be set to 10 to 20 seconds, and the second time to be within 5 seconds. The image analysis server 50 uses multimodal AI to analyze the situation at the site based on the transmitted captured images and generates text data for the first or second time. The first and second times may change depending on the situation at the site. Since the image analysis server 50 generates text data that is the basis for generating audio signals, the text data described for the first or second time refers to the time it takes to pronounce a word, phrase, or sentence when the text data is converted to speech at a normal speed.

[0023] When the image captures a suspect, the image analysis server 50 generates text data such as, "A man wearing black pants and a jacket, a hat, sunglasses, and a mask is on the run." When the image captures a suspect, the image analysis server 50 generates text data such as, "No suspect found."

[0024] The situation explanation text receiving unit 304 receives text data transmitted from the image analysis server 50 via the mobile phone network. The voice conversion unit 305 converts the text data received by the situation explanation text receiving unit 304 into an audio signal and supplies it to the wireless transmitting / receiving unit 11. The wireless transmitting unit of the wireless transmitting / receiving unit 11 wirelessly transmits the audio signal supplied by the voice conversion unit 305 to the headquarters 60.

[0025] The speaker 22 connected to the wireless transceiver unit 11 may output the audio of an audio signal transmitted from headquarters 60 and received by the wireless receiver unit of the wireless transceiver unit 11. The microphone 23 connected to the wireless transceiver unit 11 may supply an audio signal of voice spoken by the user 40 after operating a PTT (Push To Talk) switch (not shown) to the wireless transceiver unit 11, and the wireless transmitter unit may wirelessly transmit the audio signal to headquarters 60.

[0026] In the configuration shown in Figure 1, the smartphone 3 is configured with feature information transmitted from headquarters 60. However, headquarters 60 may also transmit feature information to the image analysis server 50, and the image analysis server 50 may then configure the feature information. In this case, if the captured image contains features that match the feature information, the image analysis server 50 analyzes the situation at the site based on the transmitted captured image, generates text data to explain the situation in the first time, and transmits it to the situation explanation text receiving unit 304. If the captured image does not contain features that match the feature information, the image analysis server 50 generates text data to explain the situation in the second time and transmits it to the situation explanation text receiving unit 304.

[0027] Furthermore, while the image analysis server 50 generates text data explaining the situation at the site and the voice conversion unit 305 converts the text data into a voice signal, the image analysis server 50 may also directly generate a voice signal explaining the situation at the site. In this case, if the captured image contains features that match the feature information, the image analysis server 50 should generate a voice signal explaining the situation at the site in a first time interval, and if the captured image does not contain features that match the feature information, it should generate a voice signal explaining the situation at the site in a second time interval.

[0028] In a configuration where the image analysis server 50 sends text data to the smartphone 3, there is the advantage that the amount of data sent from the image analysis server 50 to the smartphone 3 can be reduced.

[0029] Furthermore, the smartphone 3 transmits the captured images to an external image analysis server 50, which then analyzes the situation on site using multimodal AI. Alternatively, an image analysis device may be installed in the smartphone 3, and the smartphone 3 may analyze the situation on site and generate text data or audio signals explaining the situation. The image analysis device installed in the smartphone 3 may also use multimodal AI or a similar AI to explain the situation.

[0030] Smartphone 3 functions as a situational commentary voice acquisition unit that analyzes the situation at the site and acquires situational commentary voice that explains the situation. When the captured image taken in front of the user 40 located at the site contains features that match pre-set feature information, the situational commentary voice acquisition unit acquires a first situational commentary voice that explains the situation at the site in a first time based on the captured image. When the captured image does not contain features that match the feature information, the situational commentary voice acquisition unit acquires a second situational commentary voice that explains the situation at the site in a second time, which is shorter than the first time, based on the captured image.

[0031] The wireless transmission unit of the wireless transceiver unit 11 wirelessly transmits the first or second situation explanation audio acquired by the situation explanation audio acquisition unit to the headquarters 60 using the first wireless communication method. The first wireless communication method is a communication method that employs a half-duplex communication system.

[0032] In the live commentary audio transmission device 100 shown in Figure 1, the situation commentary audio acquisition unit (smartphone 3) transmits the captured image to an external image analysis server 50. The situation commentary audio acquisition unit transmits the captured image to the image analysis server 50 using a second wireless communication method different from the first wireless communication method. The second wireless communication method is wireless communication using a mobile phone line.

[0033] The situation commentary audio acquisition unit acquires first text data for generating a first situation commentary audio, which describes the situation at the scene in a first time interval, generated by the image analysis server 50 based on the captured image, when the captured image contains features that match the feature information. The situation commentary audio acquisition unit acquires second text data for generating a second situation commentary audio, which describes the situation at the scene in a second time interval, generated by the image analysis server 50 based on the captured image, when the captured image does not contain features that match the feature information.

[0034] The situation commentary audio acquisition unit converts the first text data into an audio signal to acquire the first situation commentary audio. The situation commentary audio acquisition unit converts the second text data into an audio signal to acquire the second situation commentary audio.

[0035] The processes performed by the live commentary audio transmission device 100 will be further explained using the flowchart shown in Figure 6. In Figure 6, the live commentary audio transmission device 100 (wireless transceiver unit 11) determines in step S1 whether or not it has received characteristic information of the target being tracked. If it has not received characteristic information of the target being tracked (NO), the live commentary audio transmission device 100 repeats the process in step S1. If it has received characteristic information of the target being tracked (YES), the live commentary audio transmission device 100 saves the characteristic information in the characteristic information setting unit 301 in step S2.

[0036] In step S3, the live commentary audio transmission device 100 (feature extraction unit 302) determines whether or not it has extracted features from the captured image that match the feature information. If it has extracted features from the captured image that match the feature information (YES), in step S4, the live commentary audio transmission device 100 instructs the image analysis server 50 to provide a description of the situation at the first time. If it has not extracted features from the captured image that match the feature information (NO), in step S5, the live commentary audio transmission device 100 instructs the image analysis server 50 to provide a description of the situation at the second time.

[0037] In step S6, the live commentary audio transmission device 100 (situation explanation text receiving unit 304) receives situation explanation text data from the image analysis server 50. In step S7, the live commentary audio transmission device 100 (speech conversion unit 305) converts the text data into an audio signal. In step S8, the live commentary audio transmission device 100 (wireless transceiver unit 11) wirelessly transmits the situation explanation audio signal to headquarters 60. In step S9, the live commentary audio transmission device 100 determines whether or not to terminate the tracking. If tracking is not terminated (NO), the live commentary audio transmission device 100 repeats the processes in steps S3 to S9. If tracking is terminated (YES), the live commentary audio transmission device 100 terminates the process.

[0038] In the configuration shown in Figure 3 or Figure 4, ending the tracking means, for example, turning off the power to the radio 1 or smartphone 3, or disconnecting the connection between the radio 1 and smartphone 3. In the configuration shown in Figure 5, ending the tracking means, for example, turning off the power to smartphone 3.

[0039] The method for transmitting live commentary audio according to the first embodiment is as follows: A camera 21 attached to a user 40 located at the site takes a picture of the area in front of the user 40. The situation commentary audio acquisition unit (smartphone 3) acquires a first situation commentary audio that describes the situation at the site in a first time based on the captured image if the captured image contains features that match pre-set feature information. The situation commentary audio acquisition unit (smartphone 3) acquires a second situation commentary audio that describes the situation at the site in a second time shorter than the first time based on the captured image if the captured image does not contain features that match the feature information.

[0040] The wireless transmission unit (wireless transceiver unit 11) wirelessly transmits the first or second situational commentary audio to the headquarters 60.

[0041] According to the live commentary audio transmission device and live commentary audio transmission method of the first embodiment, the situation at the site can be automatically transmitted to the headquarters 60 in audio. According to the live commentary audio transmission device and live commentary audio transmission method of the first embodiment, when the image captured by the camera 21 contains features that match pre-set feature information, a relatively long audio commentary on the situation is transmitted to the headquarters 60, and when no matching features are found, a relatively short audio commentary on the situation is transmitted to the headquarters 60. The headquarters 60 can obtain detailed audio information when the captured image contains features that match pre-set feature information, and can obtain concise audio information on the situation when no matching features are found.

[0042] <Second Embodiment> Figure 7 shows a live commentary audio transmission device 200 according to the second embodiment. In Figure 7, the same parts as in Figure 1 are denoted by the same reference numerals, and their descriptions may be omitted. As shown in Figure 7, the radio 1 has a short-range wireless communication unit 101, a priority setting unit 102, and a transmission timing control unit 103. One of the two users 40 will be referred to as user 40A, and the other as user 40B. The live commentary audio transmission device 200 installed on user 40A will be referred to as live commentary audio transmission device 200A, and the components within live commentary audio transmission device 200A will be denoted by the reference numeral A. The live commentary audio transmission device 200 installed on user 40B will be referred to as live commentary audio transmission device 200B, and the components within live commentary audio transmission device 200B will be denoted by the reference numeral B.

[0043] If radio 1 and headquarters 60 transmit and receive audio signals on a single channel, and the live commentary audio transmitters 200A and 200B transmit audio signals simultaneously, a collision will occur between the transmissions of one and the other, and headquarters 60 will not be able to receive audio signals simultaneously. Live commentary audio transmitter 200 is configured to control the timing of audio signal transmission so as to reduce the occurrence of collisions when live commentary audio transmitters 200A and 200B transmit audio signals to headquarters 60.

[0044] The live audio transmission devices 200A and 200B communicate with each other via the short-range wireless communication unit 101. As an example, the short-range wireless communication unit 101 can adopt Bluetooth (registered trademark) as the short-range wireless communication standard. Suppose the live audio transmission device 200A extracts features that match the feature information from the captured image first. The short-range wireless communication unit 101A in the live audio transmission device 200A sets the priority number No. 1 for its own live audio transmission device 200A.

[0045] The short-range wireless communication unit 101A notifies the short-range wireless communication unit 101B in the live audio transmission device 200B that the priority number No. 1 has been set for the live audio transmission device 200A. Also, the short-range wireless communication unit 101A transmits to the short-range wireless communication unit 101B the timing including the transmission start time when the live audio transmission device 200A wirelessly transmits the voice signal of the situation explanation to the headquarters 60 and the first time at that time.

[0046] The priority setting unit 102A in the live audio transmission device 200A sets that the live audio transmission device 200A has the priority number No. 1 and the live audio transmission device 200B has the priority number No. 2. Since the live audio transmission device 200A is set to have the priority number No. 1 in the priority setting unit 102A, the transmission timing control unit 103A controls the wireless transceiver unit 11A to transmit at the timing when the wireless transceiver unit 11A acquires the voice signal supplied from the voice conversion unit 305A.

[0047] The priority setting unit 102B in the live audio transmission device 200B sets that the live audio transmission device 200A has the priority number No. 1 and the live audio transmission device 200B has the priority number No. 2. The transmission timing control unit 103B controls the wireless transceiver unit 11B to transmit at the timing after the first time has elapsed from the timing including the transmission start time when the live audio transmission device 200A wirelessly transmits the voice signal of the situation explanation to the headquarters 60, when the wireless transceiver unit 11B acquires the voice signal supplied from the voice conversion unit 305B.

[0048] In this way, when the live audio transmission device 200 transmits the first or second situation explanation audio to the headquarters 60, the priority setting unit 102 sets the priority, and the transmission timing control unit 103 transmits the first or second situation explanation audio to the headquarters 60 according to the priority set by the priority setting unit 102. Therefore, even when the radio 1 and the headquarters 60 transmit the situation explanation audio on one channel, the occurrence of collisions between the live audio transmission devices 200A and 200B can be reduced.

[0049] The plurality of live audio transmission devices 200 communicate with each other by a third wireless communication method different from the first and second wireless communication methods, and the priority setting unit 102 sets the priority.

[0050] As shown in FIG. 8A, assume that the live audio transmission device 200A transmits the situation explanation audio of the time TA1 as the first time to the headquarters 60 at time t1. Assume that the live audio transmission device 200B also extracts a feature that matches the feature information from the captured image after the live audio transmission device 200A. The live audio transmission device 200B can transmit the situation explanation audio of the time TB1 as the first time to the headquarters 60 at time t2 after the transmission of the situation explanation audio by the live audio transmission device 200A is completed.

[0051] FIG. 8B shows the transmission of the situation explanation audio when the live audio transmission device 200B does not extract a feature that matches the feature information from the captured image. The live audio transmission device 200B can transmit the situation explanation audio of the time TB2 as the second time to the headquarters 60 at time t2 after the transmission of the situation explanation audio by the live audio transmission device 200A is completed.

[0052] By the way, if neither the live commentary audio transmitters 200A nor 200B extract features from the captured images, there is a possibility of a collision occurring when the live commentary audio transmitters 200A and 200B transmit the situational commentary audio for the second time period to headquarters 60. However, since the second time period is very short, the possibility of a collision occurring is low. In order to reliably avoid a collision when the live commentary audio transmitters 200A and 200B transmit the situational commentary audio for the second time period to headquarters 60, one of the live commentary audio transmitters 200A and 200B may be assigned priority No. 1 and the other priority No. 2 in advance.

[0053] The processes performed by the live commentary audio transmission device 200 will be further explained using the flowchart shown in Figure 9. In Figure 9, the same reference numerals are used for parts that are the same as those in Figure 6, and their explanations may be omitted.

[0054] In Figure 9, the live commentary audio transmitter 200 (priority setting unit 102) determines in step S11, following step S4, whether or not priority No. 1 is set for another live commentary audio transmitter 200. If priority No. 1 is not set for another live commentary audio transmitter 200 (NO), the live commentary audio transmitter 200 sets its own live commentary audio transmitter 200 as priority No. 1 in step S12. If priority No. 1 is set for another live commentary audio transmitter 200 (YES), the live commentary audio transmitter 200 sets its own live commentary audio transmitter 200 as priority No. 2 in step S13.

[0055] In step S14, following step S7, the live commentary audio transmitter 200 (transmission timing control unit 103) determines whether or not priority No. 1 is set for itself. If priority No. 1 is set for itself (YES), in step S8, the live commentary audio transmitter 200 wirelessly transmits an audio signal of the situation to headquarters 60. If priority No. 1 is not set for itself (NO), in step S15, the live commentary audio transmitter 200 determines whether or not it has received the transmission timing and transmission time from another live commentary audio transmitter 200.

[0056] If the transmission timing and transmission time have not been received from another live commentary audio transmitter 200 in step S15 (NO), the live commentary audio transmitter 200 wirelessly transmits an audio signal of situation explanation to headquarters 60 in step S8 and proceeds to step S9. If the transmission timing and transmission time have been received from another live commentary audio transmitter 200 in step S15 (YES), the live commentary audio transmitter 200 wirelessly transmits an audio signal of situation explanation to headquarters 60 in step S16 after the audio signal has been transmitted by the other live commentary audio transmitter 200 and proceeds to step S9.

[0057] According to the second embodiment of the live commentary audio transmission device and live commentary audio transmission method, in addition to the effects achieved by the first embodiment of the live commentary audio transmission device and live commentary audio transmission method, it is possible to reduce or eliminate the occurrence of collisions between multiple live commentary audio transmission devices 200 when wirelessly transmitting the situational commentary audio signal to the headquarters 60.

[0058] The present invention is not limited to the first or second embodiments described above, and can be modified in various ways without departing from the spirit of the invention. Some configurations of the live commentary audio transmission device 100 or 200 may be realized by the CPU executing a computer program. In particular, the configuration within the smartphone 3 can be realized by the CPU 34 executing a computer program.

[0059] This application claims priority based on Japanese Patent Application No. 2024-163410, filed with the Japan Patent Office on September 20, 2024, the full disclosure of which is incorporated herein by reference.

Claims

1. A live commentary audio transmission device comprising: a situation commentary audio acquisition unit that, when a captured image taken in front of a user located at a site contains features that match pre-set feature information, acquires a first situation commentary audio that explains the situation at the site in a first time based on the captured image; and when the captured image does not contain features that match the feature information, acquires a second situation commentary audio that explains the situation at the site in a second time shorter than the first time based on the captured image; and a wireless transmission unit that wirelessly transmits the first or second situation commentary audio acquired by the situation commentary audio acquisition unit to a predetermined device using a first wireless communication method.

2. The situation commentary audio acquisition unit transmits the captured image to an external image analysis server; when the captured image contains features matching the feature information, acquires first text data for generating the first situation commentary audio, which is generated by the image analysis server based on the captured image and describes the situation at the site in a first time; when the captured image does not contain features matching the feature information, acquires second text data for generating the second situation commentary audio, which is generated by the image analysis server based on the captured image and describes the situation at the site in a second time; converts the first text data into an audio signal to acquire the first situation commentary audio; and converts the second text data into an audio signal to acquire the second situation commentary audio.

3. The live commentary audio transmission device according to claim 2, wherein the predetermined device receives the first or second situational commentary audio from a plurality of the live commentary audio transmission devices, and each of the live commentary audio transmission devices comprises: a priority setting unit that sets the priority for transmitting the first or second situational commentary audio to the predetermined device; and a transmission timing control unit that transmits the first or second situational commentary audio to the predetermined device according to the priority set by the priority setting unit.

4. The situation commentary audio acquisition unit transmits the captured image to the image analysis server using a second wireless communication method different from the first wireless communication method, and the plurality of live commentary audio transmission devices communicate with each other using a third wireless communication method different from the first and second wireless communication methods, and the priority setting unit sets the priority, as described in claim 3.

5. A live commentary audio transmission method comprising: a camera attached to a user located at a site captures the area in front of the user; when the captured image from the camera contains features that match pre-set feature information, a situation commentary audio acquisition unit acquires a first situation commentary audio that describes the situation at the site in a first time based on the captured image; when the captured image does not contain features that match the feature information, the situation commentary audio acquisition unit acquires a second situation commentary audio that describes the situation at the site in a second time shorter than the first time based on the captured image; and a wireless transmission unit wirelessly transmits the first or second situation commentary audio to a predetermined device.

Citation Information

Patent Citations

  • Information processing device, information processing method, and computer program

    JP2021149309A

  • Information processing device and information processing method

    JP2022112897A

  • Wearable system for detection of environmental hazards

    US11935384B1