Voice channel configuration method, device and computer equipment
By identifying and configuring voice channels in a video surveillance system and determining priorities based on the characteristic information of intercom requests, the problem of uneven allocation of voice channel resources is solved, thereby improving the utilization rate of voice channels in the video surveillance system and the processing efficiency in emergency situations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2023-01-09
- Publication Date
- 2026-07-31
AI Technical Summary
In existing video surveillance systems, uneven allocation of audio channel resources leads to multiple users vying for the audio channel, reducing the effective utilization rate of the video surveillance system in real-time monitoring of daily operations.
By acquiring real-time monitoring video data from the video channel, preset actions are identified to determine intercom requests. Based on feature information such as facial image features, personnel distance, the frequency of occurrence of preset actions, and request time, the priority of intercom requests is determined, and the voice channel is configured according to the priority.
This effectively prevents multiple people from vying for the voice channel, improves the utilization rate of the voice channel in the video surveillance system during operation, and ensures timely handling and balanced resource allocation in emergency situations.
Smart Images

Figure CN116132619B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of surveillance technology, and in particular to a voice channel configuration method, apparatus, and computer equipment. Background Technology
[0002] Video surveillance systems are currently the technological foundation for real-time monitoring of key departments or important locations in various industries. Management departments can use them to obtain effective data, images, or audio information, monitor and record the process of sudden abnormal events in a timely manner, and provide efficient and timely command, deployment of early warning plans, and handling of emergency security incidents.
[0003] In existing video surveillance systems, when a safety incident occurs or routine work instructions and feedback are needed, staff typically rely on voice commands to initiate voice intercom requests to facilitate information exchange between front-end and back-end monitoring devices. However, in practical applications, the number of input audio channels is limited. When multiple people initiate voice intercom requests simultaneously, the aforementioned method of determining voice channel connections through voice command recognition can lead to multiple users vying for voice channels, resulting in uneven distribution of voice channel resources and ultimately reducing the effective utilization rate of the video surveillance system's voice channels in real-time monitoring of daily operations.
[0004] Currently, there is no effective solution proposed for improving the utilization rate of the voice channel in video surveillance systems during operation. Summary of the Invention
[0005] Therefore, it is necessary to provide a voice channel configuration method, apparatus, computer equipment, and storage medium that can improve the effective utilization rate of the voice channel in a video surveillance system during operation, in order to address the aforementioned technical problems.
[0006] Firstly, this application provides a method for configuring a voice channel, the method comprising:
[0007] Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel.
[0008] Based on the real-time monitoring video data, the characteristic information of each intercom request is determined;
[0009] The priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority.
[0010] In one embodiment, the feature information includes facial image features corresponding to the intercom request and the distance between people. Determining the priority of each intercom request based on the feature information includes:
[0011] Based on the facial image features, the intercom request from the same person is identified as a repeated intercom request;
[0012] Based on the distance between the people, the retained intercom requests are determined, and the remaining duplicate intercom requests are deleted.
[0013] In one embodiment, the feature information includes the number of times the preset action occurs, and determining the priority of each intercom request based on the feature information includes:
[0014] The priority of the intercom request is determined based on the number of times it occurs.
[0015] In one embodiment, determining the priority of each intercom request based on the feature information further includes:
[0016] If the number of occurrences is the same, the priority of the intercom request is determined based on the distance between the people.
[0017] In one embodiment, the feature information further includes the request time, and determining the priority of each intercom request based on the feature information further includes:
[0018] If the number of occurrences and the distance to the person are both the same, the priority of the intercom request is determined based on the request time.
[0019] In one embodiment, configuring a voice channel for each intercom request based on the priority includes:
[0020] Obtain the real-time occupancy status of each voice channel, and allocate a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority.
[0021] In one embodiment, after allocating a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority, the method further includes:
[0022] Based on the priority, a corresponding intercom queue is generated for each voice channel, and the position of the intercom request in the intercom queue is determined.
[0023] Secondly, this application also provides a voice channel configuration device, the device comprising:
[0024] The identification module is used to acquire real-time monitoring video data of the video channel. If a preset action appears in the real-time monitoring video data, it is determined that there is an intercom request in the corresponding video channel.
[0025] The feature extraction module is used to determine the feature information of each intercom request based on the real-time monitoring video data;
[0026] The determining module determines the priority of each intercom request based on the feature information, and configures a voice channel for each intercom request based on the priority.
[0027] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0028] Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel.
[0029] Based on the real-time monitoring video data, the characteristic information of each intercom request is determined;
[0030] The priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority.
[0031] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0032] Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel.
[0033] Based on the real-time monitoring video data, the characteristic information of each intercom request is determined;
[0034] The priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority.
[0035] The aforementioned voice channel configuration method, device, computer equipment, and storage medium can perform real-time detection of the acquired real-time monitoring video data of the video channel. If a preset action appears in the real-time monitoring video data, it is determined that there is an intercom request on the corresponding video channel. Then, based on the real-time monitoring video data, the characteristic information of each intercom request is determined. The priority of each intercom request is then determined based on the characteristic information, and finally, a voice channel is configured for each intercom request based on the aforementioned priority. The priority of intercom requests can be determined according to the urgency of the request and the actual request situation, avoiding uneven distribution of voice resources caused by multiple people vying for the voice channel, thereby improving the effective utilization rate of the voice channels in the video surveillance system during operation. Attached Figure Description
[0036] Figure 1 This is an application environment diagram of a voice channel configuration method in one embodiment;
[0037] Figure 2 This is a flowchart illustrating a voice channel configuration method in one embodiment;
[0038] Figure 3 This is a flowchart illustrating a preferred embodiment of a voice channel configuration method.
[0039] Figure 4 This is a flowchart illustrating the voice channel configuration method in another preferred embodiment;
[0040] Figure 5 This is a structural block diagram of a voice channel configuration device in one embodiment;
[0041] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0043] A video surveillance system consists of a real-time control system, a monitoring system, and a management information system. The real-time control system performs real-time data acquisition, processing, storage, and feedback; the monitoring system provides 24 / 7 monitoring of various control points and can switch between multiple video feeds from multiple control points; the management information system collects, receives, transmits, processes, and handles various required information. Therefore, video surveillance systems are widely used in many applications due to their intuitiveness, convenience, and rich information content. Examples include traffic safety monitoring, coal mine monitoring, and factory operation monitoring.
[0044] For example, coal mining operations are characterized by harsh environments, numerous and dispersed locations, leading to many safety hazards. Therefore, reducing mining risks and improving worker safety are critical issues that urgently need to be addressed in the coal industry.
[0045] Most coal mines have established video surveillance systems to obtain real-time operational status. During normal mine operations, information exchange between front-end and back-end monitoring devices enables work guidance and feedback. In the event of a mine accident, this information exchange allows for the acquisition of real-time status information of workers to facilitate subsequent rescue efforts. However, during this interaction, workers need to manually operate the monitoring equipment or use voice recognition technology to input and transmit information. With limited audio input channels, this leads to multiple users vying for the voice channel, resulting in uneven distribution of voice channel resources and reducing the effective utilization rate of the voice communication channels in real-time mine operation monitoring.
[0046] To address the aforementioned issues, the voice channel configuration method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed in the cloud or on other network servers. Terminal 102 first acquires real-time monitoring video data from the video channel and uploads it to server 104. Server 104 identifies the real-time monitoring data; if a preset action is detected in the real-time monitoring video data, it determines that there is an intercom request for the corresponding video channel. Then, feature extraction is performed based on the real-time monitoring video data to determine the feature information of each intercom request. Finally, the priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority. Finally, server 104 can establish an intercom connection between terminal 102 and the display device 106 in the backend. Terminal 102 can be, but is not limited to, various surveillance cameras. Server 104 can be implemented using a standalone server or a server cluster.
[0047] Figure 2 This is a flowchart illustrating a voice channel configuration method in one embodiment, such as... Figure 2 As shown, it includes the following steps:
[0048] Step S201: Obtain real-time monitoring video data of the video channel. If a preset action appears in the real-time monitoring video data, it is determined that there is an intercom request in the corresponding video channel.
[0049] The video channels can be the real-time monitoring video data corresponding to the monitoring cameras in the video surveillance system, that is, the real-time video captured by each monitoring device. Preset actions are human body movements that technicians pre-set to trigger intercom requests, including at least hand gestures and body movements. For example, preset actions could be waving an arm, raising a hand, or making an SOS gesture. Furthermore, when acquiring real-time monitoring video data, multiple video channels or a single video channel can be acquired, and the appropriate video channel can be determined based on the actual application scenario.
[0050] Specifically, intelligent detection algorithms can be used to detect pre-set actions in the acquired real-time video data. For example, if a person is seen raising their hand above their head and waving it twice or more, then the pre-set action is identified as present in the video data. If someone initiates an intercom request, the person performing the pre-set action is identified as the person making the intercom request.
[0051] Step S202: Determine the feature information of each intercom request based on the real-time monitoring video data.
[0052] The feature information consists of data used to distinguish each intercom request, including at least the distance between the intercom personnel, the number of times the preset action appears in the intercom request, and the request time. Specifically, the personnel distance is the distance between the intercom personnel and the monitoring device, the intercom personnel is the initiator of the intercom request, and the monitoring device is the device that captures the real-time monitoring video data corresponding to the intercom request. The number of times the preset action appears is the number of times the preset action appears in the corresponding real-time monitoring video data within a specified time period, and the request time is defined as the time when the intercom personnel first execute the preset action in the real-time video monitoring data.
[0053] For example, if real-time monitoring video data (Video A) is received at the current moment, and the intelligent algorithm detects that person A is continuously waving their arm in Video A, then it is determined that there is an intercom request in Video A. Person A can then be identified as the intercom person responding to the intercom request in Video A. Based on person A's position in the video, the intelligent algorithm estimates the distance from person A to the monitoring device that captured Video A to be 2 meters, and person A waves their arm 10 times within 10 seconds, with the first wave occurring at 13:00:10. Therefore, the corresponding intercom request characteristics for Video A are: person distance 2 meters, preset action occurrences 10 times, and request time 13:00:10.
[0054] Step S203: Determine the priority of each intercom request based on the feature information, and configure a voice channel for each intercom request based on the priority.
[0055] Priority can be represented by the urgency of the intercom request and the order in which it is processed. The higher the priority, the more urgent the intercom request is and the faster it needs to be processed. The lower the priority, the less urgent the intercom request is and the more it can be processed later.
[0056] For example, in a mine monitoring environment, intercom requests such as execution confirmations and process inquiries during normal operations can be classified as low-priority normal requests. Requests for assistance during safety incidents can be classified as high-priority emergency requests.
[0057] In one embodiment of this application, the priority of each intercom request can be determined based on the personnel distance, the frequency of occurrence of the preset action, and the request time obtained above. Optionally, the frequency of occurrence of the preset action can be set as the first weighting factor, the personnel distance as the second weighting factor, and the request time as the third weighting factor. When confirming the priority, the first weighting factor to the third weighting factor can be compared sequentially. The larger the first weighting factor, the higher the priority of the intercom request, and so on. If the first weighting factor cannot determine the priority of the intercom request, the larger the second weighting factor can be determined to be the higher the priority of the intercom request. Finally, the third weighting factor is considered to determine the priority of the intercom request.
[0058] After confirming the priority of each intercom request, its corresponding voice channel can be determined. Intercom requests with higher priority are configured to the voice channel with higher current processing capacity, while intercom requests with lower priority are configured to the voice channel with lower current processing capacity.
[0059] In the above-described voice channel configuration method, real-time monitoring video data of the video channel is first acquired and detected. If a preset action is detected in the real-time monitoring video data, it is determined that a communication request exists on the corresponding video channel. Then, feature information of each communication request is determined based on the real-time monitoring video data. Finally, the priority of each communication request is determined based on the feature information, and a voice channel is configured for each communication request based on the priority. The priority of each communication request is determined by considering the actual urgency, environmental location factors, and the time of the request. Then, a corresponding voice channel is allocated to each communication request based on its priority, and its position in the corresponding communication queue is determined. This avoids the inability to report in a timely manner when multiple voice channels are simultaneously occupied during an emergency security incident, thus hindering security rescue operations. It also solves the problem of uneven distribution of voice resources caused by multiple people vying for voice channels, improving the effective utilization rate of the video surveillance system's voice channels during operation.
[0060] In one embodiment, the feature information includes the facial image features corresponding to the intercom request and the distance between the person and the intercom request. Determining the priority of each intercom request based on the feature information includes: identifying intercom requests from the same person as duplicate intercom requests based on the facial image features, determining which intercom requests to retain based on the distance between the person and the intercom request, and deleting the remaining duplicate intercom requests.
[0061] Understandably, in practical applications, multiple monitoring devices can be set up in the same area to comprehensively monitor that area. Therefore, overlapping monitoring areas will exist. If a person is currently located in this overlapping area and is using a preset channel, at least two video channels will capture their behavior, generating real-time monitoring video for uploading. This means there are duplicate intercom requests. To prevent the person using multiple voice channels, these duplicate requests need to be deleted, leaving only one intercom request.
[0062] Specifically, the facial image features of the intercom personnel are first obtained through image recognition technology, facial recognition algorithms, or human feature analysis. Then, the facial image features of the personnel in different intercom requests are compared. If intercom requests with the same facial image features are detected, the intercom requests are determined to be duplicate intercom requests. At this point, the personnel distances of the multiple duplicate intercom requests need to be obtained, the intercom request with the closest personnel distance is retained, and stored in the corresponding request cache pool. In the request cache pool, the intercom requests are marked according to the corresponding video channel number, and the remaining duplicate intercom requests also need to be deleted.
[0063] The facial image features identified can include: facial features, facial shape features, and headwear features, such as glasses, hats, etc.
[0064] In an exemplary embodiment, monitoring devices A, B, and C are installed within mine area A. A monitoring area m exists within mine area A and falls within the monitoring range of these three devices. If a person m is performing a preset action within monitoring area m, and the distances of person m to monitoring devices A, B, and C are 3m, 5m, and 4m respectively, then intercom requests can be detected in all three real-time monitoring video feeds uploaded by monitoring devices A, B, and C. Furthermore, by performing facial recognition and human feature analysis on the individuals making these three intercom requests, it can be determined that these three intercom requests are duplicate requests. Therefore, two of the intercom requests need to be deleted, and one request needs to be retained. Specifically, when determining which intercom request to retain, the request corresponding to monitoring device A can be retained based on the distance to the person, while the remaining intercom requests are deleted.
[0065] In this embodiment, it is determined whether there are duplicate request events among the multiple received intercom requests, thus preventing the same person from occupying multiple voice channels simultaneously, which would lead to uneven distribution of voice channel resources. This ensures that the same person can only occupy one voice channel, thereby ensuring a uniform distribution of voice channel resources.
[0066] In one embodiment, the feature information includes the number of times the preset action occurs, and determining the priority of each intercom request based on the feature information includes: determining the priority of the intercom request based on the number of occurrences.
[0067] The frequency of a preset action is defined as the number of times a preset action occurs within a specified duration in the real-time monitoring video data. The specified duration can be determined based on the possible situations that may occur in the actual scene.
[0068] Optionally, when determining the priority of intercom requests, intercom requests that occur twice or more can be classified as emergency intercoms, and intercom requests that occur once can be defined as normal intercoms. Emergency intercoms have a higher priority than normal intercoms. Furthermore, for emergency intercoms, the priority of different emergency intercoms can also be determined according to the frequency of occurrence of preset actions, with the emergency intercom request that occurs most frequently being determined as the highest priority intercom request, and then the priority of the corresponding intercom request decreasing as the frequency of occurrence decreases.
[0069] Preferably, when determining the number of times a preset action occurs in an intercom request, the specific number of occurrences can be determined by determining the pause time between the preset actions. For example, in real-time monitoring video data, person A is raising both hands and waving. A left-right back-and-forth wave is defined as a preset action. The number of times person A waves left and right and the pause time between each wave need to be counted. If each pause time is less than 1 second, and the wave is performed 5 times within 15 seconds, then the number of occurrences is determined to be 5. If person A also waves 5 times within 15 seconds, but the pause time between the second and third wave is greater than 1 second but less than 3 seconds, then it is determined that person A initiated two intercom requests within those 15 seconds, with the two intercom requests occurring 2 and 3 times respectively, and these two intercom requests are determined to be consecutive events.
[0070] Optionally, adjacent intercom requests within the same video channel with a time difference of less than 2 seconds can be defined as consecutive events. The algorithm in the front-end camera can first determine whether a person-initiated intercom is valid, and then the back-end system can count the number of times a preset action appears in valid requests within a certain time period. The criterion for determining a valid request can be whether the person performs a preset action.
[0071] In this embodiment, the frequency of preset actions in intercom requests is statistically analyzed, and the priority of intercom requests is determined based on the number of occurrences. The urgency of intercom requests with different frequencies is considered, ensuring that intercom requests with higher urgency have higher priority. This facilitates the priority processing of more urgent intercom requests during subsequent voice channel allocation, preventing situations where emergency events cannot be reported to the backend. For example, in a mine video monitoring system, intercom requests related to safety events can be uploaded promptly and prioritized for processing, improving the timeliness and success rate of safety rescues.
[0072] In one embodiment, the feature information further includes the request time. Determining the priority of each intercom request based on the feature information further includes: if the number of occurrences is the same, determining the priority of the intercom request based on the request time.
[0073] Optionally, in practical application scenarios, two different intercom requests with the same number of preset actions may be encountered. Considering the above situation, this embodiment also incorporates the request time of the intercom request as the basis for determining the priority of the intercom request, and determines that the intercom request with an earlier request time has a higher priority.
[0074] Preferably, when determining the priority of intercom requests based on request time, it is also necessary to determine the request time difference between different intercom requests. Specifically, based on the request cache pool, the request time difference ΔTc of intercom requests corresponding to any two different video channels in the request cache pool can be determined. It is then determined whether ΔTc is less than 2 seconds. If it is less, the two are considered synchronously triggered intercom requests, meaning the request times of the two corresponding intercom requests are the same, and their priorities cannot be determined based on their request times. If ΔTc is not less than 2 seconds, the two are considered non-simultaneous intercom requests, and their priorities can be determined based on the order of their request times.
[0075] In this embodiment, the priority is determined according to the request time of the intercom request to ensure that, in the case that the preset action occurs the same number of times, the intercom request with an earlier request time can be processed first, which provides data support for the subsequent determination of the corresponding voice channel by priority.
[0076] In one embodiment, determining the priority of each intercom request based on the feature information further includes: determining the priority of the intercom request based on the distance between people when it is determined that the number of occurrences and the request time are the same.
[0077] As mentioned in the previous embodiment, when the request times of intercom requests from two different video channels are the same, it is impossible to determine the corresponding priority based on the request time. Therefore, this embodiment also provides a method for determining the priority of intercom requests based on the distance between people.
[0078] Specifically, intercom requests with smaller confirmed distances between people have higher priority, while intercom requests with larger confirmed distances between people have relatively lower priority.
[0079] For example, if video channel A simultaneously receives three intercom requests (A, B, and C) at 14:00:00, and the preset action in each of the three intercom requests occurs once, then the priority can be determined based on the personnel distances corresponding to intercom requests A, B, and C. Specifically, intercom request A corresponds to a personnel distance of 3m, intercom request B to a personnel distance of 2m, and intercom request C to a personnel distance of 5m. Therefore, the priority order is: intercom request B > intercom request A > intercom request C.
[0080] In this embodiment, the priority is determined based on the distance between the person making the intercom request. This ensures that, when the number of preset actions and the request time are the same, intercom requests that are closer can be processed first, providing data support for subsequently determining the corresponding voice channel based on priority.
[0081] In one embodiment, configuring a voice channel for each intercom request based on the priority includes: obtaining the real-time occupancy status of each voice channel, and allocating a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority.
[0082] After determining the priority of each intercom request, the available voice channel for each request can be identified. Optionally, it is necessary to first determine the monitoring device corresponding to each intercom request. This can be done by identifying the monitoring device to which the video channel belongs based on the video channel number of the uploaded real-time monitoring video data, i.e., identifying the monitoring device that detected the intercom request, and designating the corresponding monitoring device as the target intercom device. It is understood that existing monitoring equipment typically has at least one voice channel for voice communication; therefore, after identifying the target intercom device, it is also necessary to determine the corresponding voice channel within that target intercom device.
[0083] Specifically, the system can obtain the real-time occupancy status of each voice channel in the target intercom device, assigning high-priority intercom requests to relatively idle voice channels and low-priority intercom requests to relatively busy voice channels. The real-time occupancy status of voice channels includes whether there is currently voice input on the channel, based on the number of requests in the corresponding intercom queue. Channels with a high number of requests are determined to be currently busy with a high real-time occupancy rate, while channels with a low number of requests are determined to be currently idle with a low real-time occupancy rate.
[0084] Optionally, the real-time occupancy status of voice channels may also include the priority of all intercom requests in the current intercom queue corresponding to the voice channel. Higher-priority intercom requests can be assigned to the voice channel currently processing a lower-priority intercom request. For example, if the intercom queues corresponding to voice channels A and B both currently contain 5 requests, and the intercom requests in voice channel A's queue are all normal intercoms, while the intercom queue in voice channel B contains urgent intercom requests that are currently being processed, then the higher-priority intercom request that needs processing can be assigned to voice channel A.
[0085] In this embodiment, a corresponding voice channel is configured for each intercom request based on its priority and the real-time occupancy of the voice channel. This allows for the reasonable and efficient allocation of voice channels, taking into account the actual situation of the intercom requests, thereby improving the allocation efficiency of the voice channels. Furthermore, the actual occupancy of each voice channel is considered to determine the allocation of voice channels, avoiding uneven distribution of voice channel resources and improving the effective utilization rate of the voice channels.
[0086] In one embodiment, after allocating a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority, the method further includes: generating an intercom queue for each voice channel based on the priority, and determining the position of the intercom request in the intercom queue.
[0087] Preferably, after allocating a voice channel for each intercom request, it is also necessary to determine the corresponding intercom queue and the position of the intercom request in the intercom queue according to its corresponding priority. Specifically, the intercom requests are sorted from high to low priority to determine the order of the intercom requests, thereby generating the intercom queue.
[0088] Furthermore, if an intercom queue already exists in the current voice channel, a new intercom queue needs to be determined by combining the priorities of each intercom request in the current queue with the priority of the newly added intercom request. For example, if the priority of the intercom request at the head of the current intercom queue is 2, and the priority of the newly added intercom request is 5, then the newly added intercom request can be inserted at the head of the intercom queue. In another example, if the priority order of all intercom requests in the current intercom queue is (5, 3, 2, 1), and the priority of the newly added intercom request is 4, then the newly added intercom request can be inserted at the second position in the intercom queue, and the priority order of all intercom requests in the new intercom queue will be (5, 4, 3, 2, 1).
[0089] Furthermore, in another example, if the priority of the intercom request currently being processed by the voice channel is 2, while the priority of the newly added intercom request is 5, then the current intercom request can be paused, the newly added intercom request can be connected, and it can be processed first.
[0090] In this embodiment, the specific position of the intercom request in the intercom queue is determined by priority. This ensures that intercom requests with higher priority are processed first, avoiding situations where urgent intercom requests cannot be transmitted in a timely manner, and improving the processing efficiency of intercom requests.
[0091] Preferably, after allocating a voice channel for each intercom request, it is also necessary to monitor the real-time occupancy of the voice channel, detect whether there is a voice signal in the voice channel, and if it is detected that there is no voice signal in the current voice channel and the duration of no signal exceeds a preset time threshold, it is determined that the voice channel has been occupied for a long time but has not been used, the current intercom request is terminated, and the intercom request located at the head of the intercom queue can be connected.
[0092] Furthermore, in another embodiment, the intercom queue of this application can also be the intercom queue of the target intercom device. When determining the voice channel for each intercom request, the corresponding voice channel number can also be recorded. When it is the turn of the intercom request in the corresponding intercom queue, the voice channel on the target intercom device can be opened according to the voice channel number. For example, if the currently executing intercom request is on voice channel A, and the head position in the intercom queue is the intercom request on voice channel B, then after the current intercom request ends, voice channel A can be closed and voice channel B can be opened to process the next intercom request.
[0093] In another embodiment, the voice channel of this application may be a built-in voice channel on a monitoring device used to capture real-time monitoring video, or it may be an external voice channel connected to the monitoring device.
[0094] Figure 3This is a flowchart illustrating a preferred embodiment of the voice channel configuration method, as shown below. Figure 3 As shown, it includes:
[0095] Step S301: Acquire real-time monitoring video data and detect it using an intelligent algorithm. If a preset action is detected in the real-time monitoring video data, it is determined that there is an intercom request event in the corresponding video channel.
[0096] Step S302: Determine the characteristic information of each intercom request event based on real-time monitoring video data.
[0097] The feature information includes at least: the facial image features corresponding to the intercom request, the distance between the person and the person, the number of times the preset action occurs, and the request time.
[0098] Step S303: Filter the multiple intercom request events obtained to determine whether the intercom request event is a duplicate intercom request event. Filter the duplicate intercom request events, retain one duplicate intercom request event, and store the filtered retained intercom request event in the event cache pool.
[0099] Specifically, based on the facial image features in the feature information, intercom request events where the requester is the same person can be identified as repeated intercom request events.
[0100] Based on the personnel distance in the feature information, the duplicate intercom request events to be retained are determined, and the remaining duplicate intercom requests are deleted.
[0101] Step S304: Determine whether each intercom request event in the event cache pool is a normal intercom event or an emergency intercom event.
[0102] Specifically, the frequency of occurrence of preset actions in the feature information can be used to determine whether an intercom request event is a normal intercom event. If the frequency is less than 2, the corresponding intercom request event is determined to be a normal intercom event; if the frequency is greater than 2, the corresponding intercom request event is determined to be an emergency intercom event.
[0103] Step S305: For ordinary intercom events, determine whether the intercom request event is a synchronous event based on the request time. If so, determine the priority of the intercom request event based on the distance between the personnel. If not, determine the priority of the intercom request event based on the request time.
[0104] Among them, synchronous events are intercom request events initiated simultaneously. Specifically, the time difference ΔTc between intercom request events corresponding to any two different video channels in the event buffer pool is determined, and it is checked whether ΔTc is less than 2 seconds. If it is less, the two are determined to be synchronous events.
[0105] Step S306: For emergency intercom events, determine whether the intercom request event is a peer event based on the number of times the preset action appears in the intercom request event. If so, determine the priority of the intercom request event based on the request time. If not, determine the priority of the intercom request event based on the number of times the preset action appears.
[0106] Among them, events of the same level are intercom request events with the same degree of urgency. Specifically, the frequency of occurrence of preset actions in different intercom request events is compared. If the frequency is the same, they are determined to be events of the same level. Preferably, the priority of emergency intercom events is always higher than the priority of ordinary intercom events.
[0107] Step S307: Obtain the real-time occupancy status of each voice channel, and allocate a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority.
[0108] Step S308: Generate a corresponding intercom queue for each voice channel based on the priority, and determine the position of the intercom request in the intercom queue.
[0109] In this embodiment, real-time monitoring video data from video channels is acquired and analyzed. When a preset action is detected in the real-time monitoring video data, a communication request is determined for the corresponding video channel. Then, feature information for each communication request is determined based on the real-time monitoring video data. Communication request events are categorized into ordinary communication events and emergency communication events based on the frequency of the preset action in the feature information. For ordinary communication events, the priority is determined based on the request time and the distance to the person. For emergency events, the priority is determined based on the frequency of the preset action and the request time. Finally, a voice channel is configured for each communication request event based on priority, and a corresponding communication queue is generated. The priority of each communication request is determined by considering the actual urgency, environmental location, and request time. Then, a corresponding voice channel is assigned to each communication request based on its priority, and its position in the corresponding communication queue is determined. This avoids the situation where multiple voice channels are simultaneously occupied during an emergency, hindering timely reporting and emergency rescue efforts. It also solves the problem of uneven distribution of voice resources caused by multiple people vying for the voice channel, and improves the effective utilization rate of the voice channel of the video surveillance system during operation.
[0110] It's understandable that in the complex environment of underground mines, the likelihood of a single person being simultaneously detected by multiple front-end camera devices is very high. However, for intercom request events initiated by a single person within the current timeframe, there's no need for event filtering or priority determination. The target intercom device can be directly determined based on the distance to the person in the intercom request event corresponding to each video channel, and then the intercom request event can be processed through the target intercom device. Specifically, Figure 4 This is a flowchart illustrating the voice channel configuration method in another preferred embodiment, as shown below. Figure 4 As shown:
[0111] Step S401: Acquire real-time monitoring video data and detect it using an intelligent algorithm. If a preset action is detected in the real-time monitoring video data, it is determined that there is an intercom request event in the corresponding video channel.
[0112] Step S402: Determine the characteristic information of each intercom request event based on real-time monitoring video data.
[0113] Step S403: Determine whether the requesters of all current intercom request events are the same based on the facial image feature information in the feature information.
[0114] Step S404: If yes, then determine the target intercom device based on the personnel distance in the feature information, and process the intercom request event based on the target intercom device.
[0115] In this embodiment, when only a single person requests an intercom, the corresponding target intercom device can be directly determined based on the distance between the person and the intercom request can be processed directly through the target intercom device. This takes into account the actual intercom situation and omits some determination processes, thereby improving the allocation efficiency of intercom resources. This allows for the rapid and efficient establishment of intercom connections when intercom resources are abundant, thus improving the efficiency of intercom work.
[0116] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0117] Based on the same inventive concept, this application also provides a voice channel configuration device for implementing the voice channel configuration method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more voice channel configuration device embodiments provided below can be found in the limitations of the voice channel configuration method described above, and will not be repeated here.
[0118] In one embodiment, such as Figure 5 As shown, a voice channel configuration device is provided, including: a recognition module 51, a feature extraction module 52, and a determination module 53, wherein:
[0119] The identification module 51 is used to acquire real-time monitoring video data of the video channel. If a preset action appears in the real-time monitoring video data, it is determined that there is an intercom request in the corresponding video channel.
[0120] Feature extraction module 52 is used to determine the feature information of each intercom request based on the real-time monitoring video data;
[0121] The determining module 53 determines the priority of each intercom request based on the feature information, and configures a voice channel for each intercom request based on the priority.
[0122] In the aforementioned device, real-time monitoring video data from the video channel is first acquired and detected. If a preset action is detected in the real-time monitoring video data, it is determined that a communication request exists on the corresponding video channel. Then, based on the real-time monitoring video data, characteristic information of each communication request is determined. Finally, based on the characteristic information, the priority of each communication request is determined, and a voice channel is configured for each communication request based on the priority. The priority of each communication request is determined by considering the actual urgency of the situation, environmental location factors, and the time of the request. Then, a corresponding voice channel is allocated to each communication request based on its priority, and its position in the corresponding communication queue is determined. This avoids the inability to report in a timely manner when multiple voice channels are simultaneously occupied during an emergency security incident, thus hindering security rescue operations. It also solves the problem of uneven distribution of voice resources caused by multiple people vying for voice channels, improving the effective utilization rate of the video surveillance system's voice channels during operation.
[0123] Furthermore, the feature information includes the facial image features corresponding to the intercom request and the distance between the people. The determining module 53 is also used to determine the intercom request where the requesting person is the same person as a duplicate intercom request based on the facial image features, determine the intercom requests to be retained based on the distance between the people, and delete the remaining duplicate intercom requests.
[0124] Furthermore, the feature information includes the number of times the preset action occurs, and the determining module 53 is also used to determine the priority of the intercom request based on the number of occurrences.
[0125] Furthermore, the determining module 53 is also used to determine the priority of the intercom request based on the distance between the personnel when the number of occurrences is the same.
[0126] Furthermore, the feature information also includes the request time, and the determining module 53 is further used to determine the priority of the intercom request based on the request time when it is determined that the number of occurrences and the distance to the person are the same.
[0127] Furthermore, the determining module 53 is also used to obtain the real-time occupancy status of each voice channel, and allocate a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority.
[0128] Furthermore, the determining module 53 is also used to generate a corresponding intercom queue for each voice channel based on the priority, and to determine the position of the intercom request in the intercom queue.
[0129] Each module in the aforementioned voice channel configuration device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0130] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a voice channel configuration method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0131] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0132] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0133] Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel.
[0134] Based on the real-time monitoring video data, the characteristic information of each intercom request is determined;
[0135] The priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority.
[0136] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0137] Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel.
[0138] Based on the real-time monitoring video data, the characteristic information of each intercom request is determined;
[0139] The priority of each intercom request is determined based on the aforementioned feature information, and a voice channel is configured for each intercom request based on the priority. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0141] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0142] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A voice channel configuration method, characterized by, The method includes: Acquire real-time monitoring video data of the video channel; if a preset action appears in the real-time monitoring video data, determine that there is an intercom request in the corresponding video channel. The feature information of each intercom request is determined based on the real-time monitoring video data; the feature information includes the distance between the intercom personnel and the personnel, the number of times the preset action occurs, and the request time of the intercom request; The priority of each intercom request is determined based on the feature information, and a voice channel is configured for each intercom request based on the priority. Determining the priority of each intercom request based on the feature information includes: The priority of each intercom request is determined by comparing the frequency of the preset action, the distance between the people, and the request time.
2. The method of claim 1, wherein, The feature information includes the facial image features corresponding to the intercom request. The priority of each intercom request is determined based on the feature information, including: Based on the facial image features, the intercom request from the same person is identified as a repeated intercom request; Based on the distance between the people, the retained intercom requests are determined, and the remaining duplicate intercom requests are deleted.
3. The method according to claim 1, characterized in that, Determining the priority of each intercom request based on the aforementioned feature information includes: The priority of the intercom request is determined based on the number of times it occurs.
4. The method according to claim 3, characterized in that, Determining the priority of each intercom request based on the aforementioned feature information further includes: If the number of occurrences is the same, the priority of the intercom request is determined based on the request time.
5. The method of claim 4, wherein, Determining the priority of each intercom request based on the feature information further includes: If the number of occurrences and the request time are both the same, the priority of the intercom request is determined based on the distance between the people.
6. The method of claim 1, wherein, The step of configuring a voice channel for each intercom request based on the priority includes: Obtain the real-time occupancy status of each voice channel, and allocate a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority.
7. The method of claim 6, wherein, After allocating a corresponding voice channel for each intercom request based on the real-time occupancy status and the priority, the method further includes: Based on the priority, a corresponding intercom queue is generated for each voice channel, and the position of the intercom request in the intercom queue is determined.
8. A voice path configuration apparatus characterized by comprising: The device includes: The identification module is used to acquire real-time monitoring video data of the video channel. If a preset action appears in the real-time monitoring video data, it is determined that there is an intercom request in the corresponding video channel. The feature extraction module is used to determine the feature information of each intercom request based on the real-time monitoring video data; the feature information includes the distance between the intercom personnel and the personnel, the number of times the preset action occurs, and the request time of the intercom request; The determining module determines the priority of each intercom request based on the feature information, and configures a voice channel for each intercom request based on the priority; the determination of the priority of each intercom request based on the feature information includes: sequentially comparing the occurrence frequency of the preset action, the distance between the personnel, and the request time to determine the priority of each intercom request.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.