Unmanned mine card and electric shovel cooperative monitoring method and device based on multi-mode voiceprint
By setting up multiple sound source monitoring points around the mining truck and electric shovel, calculating the spatial location of the sound source and screening the concentrated sound pattern segments, the problem of sound source feature extraction in the collaborative operation of unmanned mining trucks and electric shovels was solved, achieving accurate monitoring and safety assurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-03
AI Technical Summary
In mining operations, the coordinated operation of unmanned mining trucks and electric shovels is subject to environmental interference. Existing monitoring methods are easily affected by dust and noise, making it difficult to accurately identify the loading status, resulting in monitoring blind spots and misjudgments. It is also impossible to effectively extract key sound source characteristics, affecting operational safety and efficiency.
By setting up multiple sound source monitoring points around the mining truck and electric shovel, the spatial location of the sound source is calculated using the time difference and sound speed of the same sound signal received by the monitoring points. The concentrated sound pattern segment is screened out, the calibration characteristic value is calculated and the peak point is selected. The sound pattern peak value and frequency are compared with the preset stacking interval, and the verification signal is output.
It achieves precise location of sound source, eliminates environmental interference, focuses on core voiceprint information, improves voiceprint feature recognition and distinguishability, ensures the safety and stability of unmanned mining trucks and electric shovels working together, and provides real-time monitoring and anomaly early warning support.
Smart Images

Figure CN121784670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative technology of unmanned mining trucks and electric shovels, and in particular to a method, device, equipment and medium for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints. Background Technology
[0002] In mining operations, the coordinated operation of unmanned mining trucks and electric shovels is a core element in improving mining efficiency and reducing labor costs. The accuracy of this coordination directly affects operational safety and production stability. However, the mining environment is characterized by high dust levels, strong noise, and frequent vibrations. Traditional monitoring methods that rely on visual sensors (such as cameras) or single physical parameters (such as location) are easily affected by environmental interference, resulting in blind spots or risks of misjudgment. For example, visual equipment is prone to image blurring due to dust cover, making it difficult to accurately identify the loading status of the mining truck. Simple location data can only reflect the relative position of the equipment and cannot perceive key information such as the height of ore accumulation and abnormal equipment movements during the loading process in real time.
[0003] The core of collaboration between unmanned mining trucks and electric shovels lies in the efficient connection of "excavation-loading-transportation," among which real-time judgment of the stacking height of the mining trucks is crucial to avoiding overloading and optimizing loading rhythm. In existing technologies, some solutions monitor stacking height using weight sensors or lidar. However, weight sensors require modifications to the mining truck structure, resulting in high costs and susceptibility to vibration; lidar is prone to data inaccuracies due to ore splashes and dust obstruction. Furthermore, the sound source signals of mining equipment (impact sounds during unloading, mechanical operating sounds) contain rich operational status information, but traditional voiceprint technology often relies on a single microphone to collect data, making it difficult to distinguish effective sound sources in complex environments, and even more difficult to combine sound source location and characteristics to achieve accurate status inversion, resulting in the underutilization of the application value of voiceprint information. Therefore, how to effectively extract key sound source features in the collaborative operation of unmanned mining trucks and electric shovels through multi-dimensional perception and precise analysis technologies, and achieve reliable monitoring of stacking status and operational safety, has become an important issue for improving the level of mine intelligence.
[0004] Therefore, there is an urgent need for a collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal acoustic signatures to solve the technical problem of effectively extracting key sound source features in the collaborative operation of unmanned mining trucks and electric shovels, and reliably monitoring the stacking status and operational safety in multidimensional perception and precise analysis technologies. Summary of the Invention
[0005] To overcome the problems existing in related technologies, this disclosure provides a method, device, equipment and medium for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal acoustic signatures, in order to solve the technical problem of effectively extracting key sound source features in the collaborative operation of unmanned mining trucks and electric shovels, and reliably monitoring stacking status and operational safety in multi-dimensional perception and accurate analysis technologies in related technologies.
[0006] This specification provides one or more embodiments of a method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal acoustic signatures, including the following steps: Multiple sound source monitoring points are set up around the mining truck and electric shovel. The spatial location of the sound source is calculated based on the time difference and sound speed of the same sound signal received by different sound source monitoring points. Based on the spatial location of the sound source, the location of the sound source in the working area of the mining card and electric shovel is taken as the soundprint verification point. The soundprint features associated with the soundprint verification point are checked and verified to select the concentrated soundprint segment. Calculate the calibration feature value within the voiceprint concentration segment, and select the peak point corresponding to the calibration feature value as the feature point; The peak value and frequency of the voiceprint of the feature point are compared with the preset stacking interval, and the corresponding verification signal is output according to whether the proportion of the cross range exceeds the preset threshold.
[0007] Preferably, the step of calculating the spatial location of the sound source based on the time difference and speed of sound when the same sound signal is received at different sound source monitoring points specifically includes the following steps: Obtain the time when different sound source monitoring points receive the same sound waveform; Select two sets of reception times, calculate the time difference, and calculate the difference distance based on the speed of sound; Based on the difference distance and the spatial location of each monitoring point, connect the spatial locations to construct a spatial plane perpendicular to the line connecting the monitoring points; Multiple spatial planes are acquired, and the spatial location of the sound source is determined based on the intersection of the multiple spatial planes.
[0008] Preferably, the step of verifying and checking the voiceprint features associated with the voiceprint verification points and selecting the voiceprint concentration segments specifically includes the following steps: Vertical lines are set at both ends of the acoustic waveform and the two sets of vertical lines are controlled to move towards each other at varying rates to divide the waveform into multiple partial bands. Calculate the lumped characteristic values for each band; Standard bands with concentrated feature values not exceeding a preset value are selected, and the standard band with the largest concentrated feature value is taken as the voiceprint concentration segment.
[0009] Preferably, the step of calculating the calibration feature value within the voiceprint concentration segment and selecting the peak point corresponding to the calibration feature value as the feature point specifically includes the following steps: Analyze the peak points within the aforementioned voiceprint concentration segment and calculate the first feature of each peak point; The calibration feature value is calculated based on the first feature and the peak value of the voiceprint. The peak point with the largest calibration eigenvalue is selected as the feature point.
[0010] Preferably, the step of comparing the peak value and frequency of the feature points with a preset stacking interval, and outputting a corresponding verification signal based on whether the proportion of the intersection range exceeds a preset threshold, specifically includes the following steps: The soundprint peak value and frequency are compared with preset empty compartment intervals, low stacking intervals, medium stacking intervals and high stacking intervals respectively to confirm the set interval to which the soundprint peak value belongs. Each stacking interval corresponds to a different soundprint peak value range and frequency range. The frequency range associated with the set interval is used as the standard range, and the endpoint frequency of the voiceprint concentration segment is used as the verification range. Calculate the intersection range between the verification range and the standard range of each interval. If the intersection range accounts for ≥80% of the standard range, output the verification signal corresponding to the stacking state. If the proportion is <80%, re-screen the voiceprint concentration segment and repeat the comparison steps.
[0011] This specification provides one or more embodiments of a multimodal acoustic signature-based unmanned mining truck and electric shovel collaborative monitoring device, including a sound source location determination module, an acoustic signature concentration segment screening module, a feature point screening module, and a stacking status judgment module. The sound source location determination module is used to set up multiple sound source monitoring points around the mining truck and electric shovel, and calculate the spatial location of the sound source based on the time difference and sound speed of the same sound signal received by different sound source monitoring points. The soundprint concentration segment screening module is used to select the sound source location points in the working area of the mining card and electric shovel as soundprint verification points based on the spatial location of the sound source, and to verify the soundprint features associated with the soundprint verification points to screen out the soundprint concentration segments. The feature point filtering module is used to calculate the calibration feature value within the voiceprint concentration segment and select the peak point corresponding to the calibration feature value as the feature point. The stacking state judgment module is used to compare the peak value and frequency of the voiceprint of the feature point with the preset stacking interval, and output the corresponding verification signal according to whether the proportion of the cross range exceeds the preset threshold.
[0012] Preferably, the sound source location determination module is further configured to: Obtain the time when different sound source monitoring points receive the same sound waveform; Select two sets of reception times, calculate the time difference, and calculate the difference distance based on the speed of sound; Based on the difference distance and the spatial location of each monitoring point, connect the spatial locations to construct a spatial plane perpendicular to the line connecting the monitoring points; Multiple spatial planes are acquired, and the spatial location of the sound source is determined based on the intersection of the multiple spatial planes.
[0013] Preferably, the voiceprint concentration segment filtering module is further configured as follows: Vertical lines are set at both ends of the acoustic waveform and the two sets of vertical lines are controlled to move towards each other at varying rates to divide the waveform into multiple partial bands. Calculate the lumped characteristic values for each band; Standard bands with concentrated feature values not exceeding a preset value are selected, and the standard band with the largest concentrated feature value is taken as the voiceprint concentration segment.
[0014] This specification provides one or more embodiments of a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints.
[0015] This specification provides one or more embodiments of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints.
[0016] This disclosure provides a method, device, equipment, and medium for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal acoustic signatures. The advantage lies in that by setting up multiple sound source monitoring points around the mining truck and electric shovel, and calculating the spatial location of the sound source based on the time difference and sound velocity of the same sound signal received by different sound source monitoring points, the spatial location of the sound source is accurately located. This method overcomes the limitation of a single monitoring point not being able to determine spatial coordinates, ensuring that the location information of all sound sources within the working area of the mining truck and electric shovel can be quantitatively captured. Based on the spatial location of the sound sources, the location points of the sound sources within the working area of the mining truck and electric shovel are used as soundprint verification points. The soundprint features associated with the verification points are checked and verified to filter out concentrated soundprint segments. Based on the determined spatial location of the sound sources, soundprint verification points are set. By checking and verifying the soundprint features associated with the verification points, concentrated soundprint segments are filtered out, effectively eliminating non-target soundprints such as environmental interference sounds and irrelevant equipment operation sounds. This achieves preliminary purification of soundprint information, focusing scattered and messy soundprint data on the core soundprint intervals directly related to the operation of the mining truck and electric shovel, reducing redundancy in subsequent feature calculations. The process involves: Calculating the calibration feature values within the concentrated acoustic signature segment, selecting the peak points corresponding to these calibration feature values as feature points, and transforming continuous acoustic signature data into key feature nodes. This focuses on the most representative peak information within the concentrated acoustic signature segment, extracting quantitative indicators that reflect the core characteristics of the working status of the mining truck and electric shovel. This transforms ambiguous acoustic signature information into clear and comparable feature parameters, significantly improving the recognizability and distinguishability of acoustic signature features. The acoustic signature peak values and frequencies of the feature points are compared with preset stacking intervals. Based on whether the percentage of the overlapping range exceeds a preset threshold, a corresponding verification signal is output, achieving a quantitative judgment of the collaborative working status of the mining truck and electric shovel. This step, through clear threshold judgment rules, transforms the feature comparison results into intuitive and executable verification signals, enabling rapid feedback on whether the collaborative operation is within the normal range. This provides direct support for real-time monitoring and anomaly warning of equipment collaboration, effectively ensuring the safety and stability of unmanned mining truck and electric shovel collaborative operations. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1A flowchart illustrating a collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal acoustic signatures, provided for one or more embodiments of this specification; Figure 2 A schematic diagram of the structure of a multimodal acoustic signature-based collaborative monitoring device for unmanned mining trucks and electric shovels, provided for one or more embodiments of this specification; Figure 3 This is a schematic diagram of the structure of a computer device provided for one or more embodiments of this specification. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this invention.
[0020] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings.
[0021] Method Implementation Examples According to embodiments of the present invention, a method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints is provided, such as... Figure 1 The diagram shown is a flowchart illustrating the collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal voiceprints provided in this embodiment. The collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal voiceprints according to this embodiment includes the following steps: S110. Multiple sound source monitoring points are set up around the mining truck and electric shovel, and the sound characteristics associated with different sound source monitoring points are verified to lock in the same type of sound pattern. Based on the time difference and sound speed of the same sound signal received by different sound source monitoring points, the spatial location of the sound source is calculated. Specifically, the spatial location of the sound source is the location where the sound is emitted. Its location is located at the intersection between the mining truck and the electric shovel. It is the sound source information generated by the material unloading process. Therefore, this type of sound pattern needs to be verified to identify the specific location of its sound source. The location of the sound pattern monitoring points of different sound source monitoring points is marked in advance by relevant personnel to facilitate the confirmation of the sound source location.
[0022] S120. In order to effectively identify whether the location of the corresponding sound is in the area between the mining truck and the electric shovel, waveform verification and location confirmation are required to ensure the accuracy of the location confirmation process.
[0023] Based on the spatial location of the sound source, the location of the sound source within the working area of the mining truck and the electric shovel is used as the soundprint verification point. The soundprint features associated with the soundprint verification point are checked and verified to screen out the soundprint concentration segment. Specifically, the soundprint concentration segment is the waveform segment with the most obvious corresponding soundprint. Within the confirmed soundprint concentration segment, the soundprint features can be effectively confirmed to identify the stacking height of the corresponding mining truck during loading, thereby ensuring that there will be no overflow or other issues during loading.
[0024] S130. Based on the confirmed voiceprint concentration segment, calculate the calibration feature value in the voiceprint concentration segment according to the variation characteristics of the peak segment in the voiceprint concentration segment, and select the peak point corresponding to the calibration feature value as the feature point.
[0025] S140. The peak value and frequency of the soundprint of the feature point are compared with the preset stacking interval. According to whether the cross-range ratio exceeds the preset threshold, the corresponding verification signal is output. Specifically, in the process of comparison and verification, there are several different standard bands, which can confirm the corresponding soundprint concentration segment. Then, the interval range is compared from the confirmed soundprint concentration segment. In the comparison process, the cross-range of the corresponding frequency can be confirmed to confirm the specific interval signal and display it for relevant personnel to view and take timely countermeasures. If the verification signal displayed on the corresponding display terminal by external personnel is a high stacking verification signal, the feeding process is stopped in time to avoid material overflow, so as to ensure the specific feeding process and effectively verify the coordination between the mining truck and the electric shovel.
[0026] The method provided in this embodiment calculates the spatial location of the sound source by setting up multiple sound source monitoring points around the mining truck and electric shovel, and using the time difference and sound speed of the same sound signal received by different sound source monitoring points. By setting up multiple sound source monitoring points around the mining truck and electric shovel, and using the time difference of the same sound signal received by multiple monitoring points in combination with the sound speed to calculate the spatial location, the precise location of the sound source is achieved. This method overcomes the limitation of a single monitoring point not being able to determine spatial coordinates, ensuring that the location information of all sound sources within the working area of the mining truck and electric shovel can be quantitatively captured. Based on the spatial location of the sound sources, the location points of the sound sources within the working area of the mining truck and electric shovel are used as soundprint verification points. The soundprint features associated with the verification points are checked and verified to filter out concentrated soundprint segments. Based on the determined spatial location of the sound sources, soundprint verification points are set. By checking and verifying the soundprint features associated with the verification points, concentrated soundprint segments are filtered out, effectively eliminating non-target soundprints such as environmental interference sounds and irrelevant equipment operation sounds. This achieves preliminary purification of soundprint information, focusing scattered and messy soundprint data on the core soundprint intervals directly related to the operation of the mining truck and electric shovel, reducing redundancy in subsequent feature calculations. The process involves: Calculating the calibration feature values within the concentrated acoustic signature segment, selecting the peak points corresponding to these calibration feature values as feature points, and transforming continuous acoustic signature data into key feature nodes. This focuses on the most representative peak information within the concentrated acoustic signature segment, extracting quantitative indicators that reflect the core characteristics of the working status of the mining truck and electric shovel. This transforms ambiguous acoustic signature information into clear and comparable feature parameters, significantly improving the recognizability and distinguishability of acoustic signature features. The acoustic signature peak values and frequencies of the feature points are compared with preset stacking intervals. Based on whether the percentage of the overlapping range exceeds a preset threshold, a corresponding verification signal is output, achieving a quantitative judgment of the collaborative working status of the mining truck and electric shovel. This step, through clear threshold judgment rules, transforms the feature comparison results into intuitive and executable verification signals, enabling rapid feedback on whether the collaborative operation is within the normal range. This provides direct support for real-time monitoring and anomaly warning of equipment collaboration, effectively ensuring the safety and stability of unmanned mining truck and electric shovel collaborative operations.
[0027] In one embodiment, the spatial location of a sound source is calculated based on the time difference and speed of sound when the same sound signal is received at different sound source monitoring points. This specifically includes the following steps: The sound waveforms associated with the electric shovel's unloading stage were confirmed, and the sound waveforms associated with different sound source monitoring points were compared to identify overlapping waveforms, i.e., waveforms associated with the same set of sound signals. The times when the overlapping waveforms—the same sound waveforms—were received by different sound source monitoring points were obtained. T i ,in i These represent different sound source monitoring points.
[0028] Two sets of receiving times were randomly selected. T i Calculate the time difference associated with the two sets of receiving times. Cs Calculate the difference distance based on the speed of sound L, L=v*Cs ,in, v is the speed of sound, representing the speed at which sound travels through the air, and is generally taken as 343 m / s.
[0029] Based on the difference distance and the spatial location of each monitoring point, the spatial locations are connected to form a connecting line, based on the difference distance. L Divide the connected lines into two groups of segments, the difference in length between the two groups of segments being equal to L Two sets of receiving times T i The line segments associated with earlier times are shorter, while those associated with later times are longer. This is because during velocity propagation, the longer the distance, the longer the propagation time, and the shorter the distance, the shorter the propagation time. Based on the location of the dividing points, a spatial plane is constructed that connects the lines perpendicular to the monitoring points.
[0030] The same processing method is applied to other receiving times in sequence to obtain multiple spatial planes associated with different sound source monitoring points, obtain the intersection points generated by multiple spatial planes during the intersection process, and determine the sound source of the overlapping waveform, i.e. the spatial location of the sound source, based on the intersection points of multiple spatial planes.
[0031] The method provided in this embodiment obtains the time when the monitoring point receives the sound, converts the time difference into a difference distance, and then constructs a spatial plane and uses multi-plane cross constraints to achieve high-precision, blind-zone-free positioning of the sound source spatial location, providing reliable positional support for subsequent voiceprint recognition.
[0032] In one embodiment, the voiceprint features associated with the voiceprint verification points are checked and verified to filter out concentrated voiceprint segments. This specifically includes the following steps: Based on the confirmed locations of different sound sources, the locations of the sounds within the working areas of the mining truck and electric shovel are recorded as soundprint verification points, and then the soundprint waveforms generated by the soundprint verification points are checked and verified.
[0033] A set of vertical lines is set at both ends of the acoustic waveform, and the two sets of vertical lines are controlled to move towards each other at a variable rate. The partial bands associated with the two sets of vertical lines are recorded in different processes. The partial bands associated with different processes are different, and multiple partial bands are obtained.
[0034] Identify the clustered features associated with different bands, confirm the peak points within each band, noting that the waveform segment preceding each peak point trends upward and the waveform segment following it trends downward. Record the total number of confirmed peak points as follows: GThen, the straight-line distance between the perpendicular lines on both sides of a certain band is denoted as... ZL Calculate the lumped characteristic values of each band. Md , Md=G / ZL .
[0035] The corresponding frequency bands associated with different motion processes are confirmed sequentially, and the concentrated features associated with these frequency bands are confirmed simultaneously. Md k ,in, k Representing different band segments, selecting concentrated characteristic values Md k Not greater than the preset value Y1 Standard band, Y1 The specific value is determined by the operator based on experience; otherwise, no standard band calibration is performed, and the maximum concentrated characteristic value is selected from the confirmed standard bands. Md k max The standard band is used as the acoustic signature concentration segment. The acoustic signature features associated with this state are the most obvious, which can achieve better data analysis results. The acoustic signature concentration segment is the specific band in the corresponding acoustic signature waveform where the feature changes are more obvious. In the subsequent analysis process, it can achieve better band feature confirmation effect, which is convenient for feature verification and comparison to confirm the status of the corresponding mine card in the loading process.
[0036] The method provided in this embodiment segments the waveform by moving vertically at both ends at varying speeds, calculates the concentrated feature values of the bands, filters standard bands, and locks the band with the largest concentrated feature value, thus accurately selecting the concentrated segment of the voiceprint. This method comprehensively covers the key areas of the voiceprint waveform and eliminates interfering bands by quantifying feature values, efficiently refining the core voiceprint information and providing high-quality data support for subsequent feature extraction.
[0037] In one embodiment, calculating the calibration feature value within the voiceprint concentration segment and selecting the peak point corresponding to the calibration feature value as the feature point specifically includes the following steps: Analyze the preceding and following points associated with different peak points within the aforementioned voiceprint concentration segment, and denot the voiceprint peak associated with the preceding point as... F1 q The peak value of the voiceprint associated with the peak point is denoted as B2q The peak value of the voiceprint associated with the later point is denoted as F2q ,in, q Representing different peak points, calculate the first feature of each peak point. D1 q , D1 q =|F1 q -B2 q|+| F2 q -B2 q | ; Based on the first feature and the peak value of the voiceprint, according to the voiceprint peak value associated with this peak point. B2 q Calculate the calibration characteristic value TZ q , TZ q =B2 q ×C1+D1 q ×C2 C1 and C2 are preset fixed coefficient factors, the specific values of which are determined by the operator based on experience. C1 is generally set to 0.684, and C2 is generally set to 0.316, to obtain different calibration characteristics associated with different peak points. TZ q .
[0038] Different calibration features associated with different peak points obtained TZ q Select the largest calibration eigenvalue TZ q max The peak points are used as feature points.
[0039] The method provided in this embodiment analyzes the first feature of the peak point in the concentrated segment of the voiceprint, calculates the calibration feature value in combination with the voiceprint peak value, and selects the peak point with the largest calibration feature value as the feature point. This accurately locks the core feature and eliminates secondary peak interference, providing a highly recognizable and reliable key feature basis for subsequent voiceprint comparison.
[0040] In one embodiment, the peak value and frequency of the voiceprint of the feature points are compared with a preset stacking interval, and a corresponding verification signal is output based on whether the proportion of the intersection range exceeds a preset threshold. Specifically, this includes the following steps: The system identifies the voiceprint peak value and its frequency associated with the feature point, compares and verifies the voiceprint peak value and frequency with a set interval, and obtains the set interval to which the voiceprint peak value belongs and the associated frequency range. The set interval includes an empty compartment interval, a low stacking interval, a medium stacking interval, and a high stacking interval. Each stacking interval corresponds to a different voiceprint peak value range and frequency range. That is, the voiceprint peak value and frequency are compared with the preset empty compartment interval, low stacking interval, medium stacking interval, and high stacking interval respectively to confirm the set interval to which the voiceprint peak value belongs and the associated frequency range.
[0041] The frequency range of the set interval to which the voiceprint peak belongs is used as the standard range, and the endpoint frequency of the voiceprint concentration segment is used as the verification range. The intersection range between the verification range and the standard range of each interval is calculated. If the proportion of the intersection range to the standard range is ≥80%, the verification signal corresponding to the stacking state is output. If the proportion is <80%, the voiceprint concentration segment is removed. The voiceprint concentration segments are re-selected from several standard bands and the comparison steps are repeated. This process is repeated to confirm and output the associated verification signal.
[0042] The empty compartment range is associated with the following numerical ranges: peak value (110dB, 200dB) and frequency range (2000Hz, 8000Hz), and is associated with the empty compartment verification signal; the low stacking range is associated with the following numerical ranges: peak value (95dB, 105dB) and frequency range (500Hz, 2000Hz), and is associated with the low stacking verification signal; the medium stacking range is associated with the following numerical ranges: peak value (85dB, 95dB) and frequency range (300Hz, 1500Hz), and is associated with the medium stacking verification signal; the high stacking range is associated with the following numerical ranges: peak value (75dB, 85dB) and frequency range (100Hz, 800Hz), and is associated with the high stacking verification signal.
[0043] The method provided in this embodiment accurately divides the stacking status into four categories: empty compartments, low / medium / high stacking, and compares the peak value and frequency of the acoustic signature. It calculates the cross-ratio using the interval frequency as the standard and the endpoint frequency as the verification range. Combined with the 80% threshold judgment and the non-compliance re-screening mechanism, it achieves accurate identification and reliable verification of the stacking status, effectively avoids misjudgment, and provides solid support for real-time status monitoring of the collaborative operation of mining trucks and electric shovels.
[0044] Device Examples According to embodiments of the present invention, a collaborative monitoring device for unmanned mining trucks and electric shovels based on multimodal voiceprints is provided, such as... Figure 2 The diagram shown is a structural schematic of the unmanned mining truck and electric shovel collaborative monitoring device based on multimodal voiceprint provided in this embodiment. The unmanned mining truck and electric shovel collaborative monitoring device based on multimodal voiceprint according to this embodiment includes a sound source location determination module 21, a voiceprint concentration segment screening module 22, a feature point screening module 23, and a stacking status judgment module 24.
[0045] The sound source location determination module 21 is used to set up multiple sound source monitoring points around the mining truck and electric shovel, and calculate the spatial location of the sound source based on the time difference and sound speed of the same sound signal received by different sound source monitoring points.
[0046] The voiceprint concentration segment screening module 22 is used to select the location of the sound source in the working area of the mining card and electric shovel as the voiceprint verification point based on the spatial location of the sound source, and to verify the voiceprint features associated with the voiceprint verification point to screen out the voiceprint concentration segment.
[0047] The feature point filtering module 23 is used to calculate the calibration feature value within the voiceprint concentration segment and select the peak point corresponding to the calibration feature value as the feature point.
[0048] The stacking state judgment module 24 is used to compare the voiceprint peak value and frequency of the feature points with the preset stacking interval, and output the corresponding verification signal according to whether the cross-range ratio exceeds the preset threshold.
[0049] The device provided in this embodiment sets up multiple sound source monitoring points around the mine car and electric shovel through the sound source location determination module 21. Based on the time difference and sound speed of the same sound signal received by different sound source monitoring points, the spatial location of the sound source is calculated. By setting up multiple sound source monitoring points around the mine car and electric shovel, and using the time difference of the same sound signal received by multiple monitoring points in combination with the sound speed to calculate the spatial location, the precise location of the sound source is achieved. This method breaks through the limitation that a single monitoring point cannot determine spatial coordinates, ensuring that the location information of all sound sources within the working area of the mining truck and electric shovel can be quantitatively captured. The soundprint concentration segment filtering module 22, based on the spatial location of the sound sources, uses the location points of the sound sources within the working area of the mining truck and electric shovel as soundprint verification points. It verifies the soundprint features associated with these verification points, filtering out concentrated soundprint segments. By setting soundprint verification points based on the determined spatial locations of the sound sources and verifying the soundprint features associated with these verification points, concentrated soundprint segments are filtered out. This effectively eliminates non-target soundprints such as environmental interference sounds and irrelevant equipment operation sounds, achieving preliminary purification of soundprint information. It focuses scattered and chaotic soundprint data onto the core soundprint intervals directly related to the operation of the mining truck and electric shovel, reducing the redundancy in subsequent feature calculations. The point filtering module 23 calculates the calibration feature values within the concentrated acoustic pattern segment, selects the peak points corresponding to the calibration feature values as feature points, and completes the transformation from continuous acoustic pattern data to key feature nodes by calculating the calibration feature values within the concentrated acoustic pattern segment and selecting peak points as feature points. It focuses on the most representative peak information within the concentrated acoustic pattern segment, extracting quantitative indicators that reflect the core characteristics of the working status of the mining truck and electric shovel, transforming ambiguous acoustic pattern information into clear and comparable feature parameters, significantly improving the recognition and distinguishability of acoustic pattern features. The stacking status judgment module 24 compares the acoustic pattern peak value and frequency of the feature points with a preset stacking interval, and outputs a corresponding verification signal based on whether the cross-range ratio exceeds a preset threshold, realizing a quantitative judgment of the collaborative working status of the mining truck and electric shovel. This step, through clear threshold judgment rules, transforms the feature comparison results into intuitive and executable verification signals, enabling rapid feedback on whether the collaborative operation is within the normal range, providing direct support for real-time monitoring and anomaly warning of equipment collaboration, and effectively ensuring the safety and stability of the collaborative operation of unmanned mining trucks and electric shovels.
[0050] In one embodiment, the sound source location determination module 21 is further configured to: Obtain the time when different sound source monitoring points receive the same sound waveform.
[0051] Select two sets of receiving times, calculate the time difference, and calculate the difference distance based on the speed of sound.
[0052] Based on the difference distance and the spatial location of each monitoring point, the spatial locations are connected to form a spatial plane perpendicular to the line connecting the monitoring points.
[0053] Multiple spatial planes are acquired, and the spatial location of the sound source is determined based on the intersection of the multiple spatial planes.
[0054] The device provided in this embodiment obtains the time when the monitoring point receives the sound, converts the time difference into a difference distance, and then constructs a spatial plane and uses multi-plane cross constraints to achieve high-precision, blind-zone-free positioning of the sound source spatial location, providing reliable positional support for subsequent voiceprint recognition.
[0055] In one embodiment, the voiceprint focusing segment filtering module 22 is further configured as follows: Vertical lines are set at both ends of the acoustic waveform, and the two sets of vertical lines are controlled to move towards each other at varying rates to divide the waveform into multiple partial bands.
[0056] Calculate the lumped characteristic values for each band.
[0057] Standard bands with concentrated feature values not exceeding a preset value are selected, and the standard band with the largest concentrated feature value is taken as the voiceprint concentration segment.
[0058] The device provided in this embodiment segments the waveform by moving vertically at both ends at varying speeds, calculates the concentrated feature values of the bands, filters standard bands, and locks the band with the largest concentrated feature value, thus accurately identifying the concentrated segment of the voiceprint. It comprehensively covers the key areas of the voiceprint waveform and eliminates interfering bands by quantifying feature values, efficiently refining the core voiceprint information and providing high-quality data support for subsequent feature extraction.
[0059] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operations of each module processing step can be understood with reference to the description of the method embodiments, and will not be repeated here.
[0060] like Figure 3 As shown, the present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints in the above embodiments, or when the computer program is executed by a processor, it implements the method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints in the above embodiments.
[0061] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0062] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are known to those skilled in the art.
Claims
1. A method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints, characterized in that, Includes the following steps: Multiple sound source monitoring points are set up around the mining truck and electric shovel. The spatial location of the sound source is calculated based on the time difference and sound speed of the same sound signal received by different sound source monitoring points. Based on the spatial location of the sound source, the location of the sound source in the working area of the mining card and electric shovel is taken as the soundprint verification point. The soundprint features associated with the soundprint verification point are checked and verified to select the concentrated soundprint segment. Calculate the calibration feature value within the voiceprint concentration segment, and select the peak point corresponding to the calibration feature value as the feature point; The peak value and frequency of the voiceprint of the feature point are compared with the preset stacking interval, and the corresponding verification signal is output according to whether the proportion of the cross range exceeds the preset threshold.
2. The method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints as described in claim 1, characterized in that, The method of calculating the spatial location of a sound source based on the time difference and speed of sound when the same sound signal is received at different sound source monitoring points includes the following steps: Obtain the time when different sound source monitoring points receive the same sound waveform; Select two sets of receiving times, calculate the time difference, and calculate the difference distance based on the speed of sound; Based on the difference distance and the spatial location of each monitoring point, connect the spatial locations to construct a spatial plane perpendicular to the line connecting the monitoring points; Multiple spatial planes are acquired, and the spatial location of the sound source is determined based on the intersection of these multiple spatial planes.
3. The method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints as described in claim 1, characterized in that, The process of verifying and checking the voiceprint features associated with the voiceprint verification points and selecting concentrated voiceprint segments includes the following steps: Vertical lines are set at both ends of the acoustic waveform and the two sets of vertical lines are controlled to move towards each other at varying rates to divide the waveform into multiple partial bands. Calculate the lumped characteristic values for each band; Standard bands with concentrated feature values not exceeding a preset value are selected, and the standard band with the largest concentrated feature value is taken as the voiceprint concentration segment.
4. The method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints as described in claim 1, characterized in that, The calculation of the calibration feature value within the voiceprint concentration segment, and the selection of the peak point corresponding to the calibration feature value as the feature point, specifically includes the following steps: Analyze the peak points within the aforementioned voiceprint concentration segment and calculate the first feature of each peak point; The calibration feature value is calculated based on the first feature and the peak value of the voiceprint. The peak point with the largest calibration eigenvalue is selected as the feature point.
5. The method for collaborative monitoring of unmanned mining trucks and electric shovels based on multimodal voiceprints as described in claim 1, characterized in that, The step of comparing the peak value and frequency of the feature points with a preset stacking interval, and outputting a corresponding verification signal based on whether the percentage of the overlap exceeds a preset threshold, specifically includes the following steps: The soundprint peak value and frequency are compared and verified with preset empty compartment intervals, low stacking intervals, medium stacking intervals and high stacking intervals respectively to confirm the set interval to which the soundprint peak value belongs. Each stacking interval corresponds to a different soundprint peak value range and frequency range. The frequency range associated with the set interval to which the voiceprint peak belongs is used as the standard range, and the endpoint frequency of the voiceprint concentration segment is used as the verification range. Calculate the intersection range between the verification range and the standard range of each interval. If the intersection range accounts for ≥80% of the standard range, output the verification signal corresponding to the stacking state. If the proportion is <80%, re-screen the voiceprint concentration segment and repeat the comparison steps.
6. A collaborative monitoring device for unmanned mining trucks and electric shovels based on multimodal voiceprints, characterized in that, It includes a sound source location determination module, a soundprint concentration segment filtering module, a feature point filtering module, and a stacking status judgment module; The sound source location determination module is used to set up multiple sound source monitoring points around the mining truck and electric shovel, and calculate the spatial location of the sound source based on the time difference and sound speed of the same sound signal received by different sound source monitoring points. The soundprint concentration segment screening module is used to select the sound source location points in the working area of the mining card and electric shovel as soundprint verification points based on the spatial location of the sound source, and to verify the soundprint features associated with the soundprint verification points to screen out the soundprint concentration segments. The feature point filtering module is used to calculate the calibration feature value within the voiceprint concentration segment and select the peak point corresponding to the calibration feature value as the feature point. The stacking state judgment module is used to compare the peak value and frequency of the voiceprint of the feature point with the preset stacking interval, and output the corresponding verification signal according to whether the proportion of the cross range exceeds the preset threshold.
7. The multimodal voiceprint-based unmanned mining truck and electric shovel collaborative monitoring device as described in claim 6, characterized in that, The sound source location determination module is further configured to: Obtain the time when different sound source monitoring points receive the same sound waveform; Select two sets of receiving times, calculate the time difference, and calculate the difference distance based on the speed of sound; Based on the difference distance and the spatial location of each monitoring point, connect the spatial locations to construct a spatial plane perpendicular to the line connecting the monitoring points; Multiple spatial planes are acquired, and the spatial location of the sound source is determined based on the intersection of these multiple spatial planes.
8. The multimodal voiceprint-based unmanned mining truck and electric shovel collaborative monitoring device as described in claim 6, characterized in that, The voiceprint concentration segment filtering module is further configured as follows: Vertical lines are set at both ends of the acoustic waveform and the two sets of vertical lines are controlled to move towards each other at varying rates to divide the waveform into multiple partial bands. Calculate the lumped characteristic values for each band; Standard bands with concentrated feature values not exceeding a preset value are selected, and the standard band with the largest concentrated feature value is taken as the voiceprint concentration segment.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal voiceprints as described in any one of claims 1 to 5.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the collaborative monitoring method for unmanned mining trucks and electric shovels based on multimodal voiceprints as described in any one of claims 1 to 5.