Train water tank water level positioning method, device, equipment and storage medium
Through pre-established sound event detection model and deep learning technology, combined with beamforming and noise reduction algorithms, the problem of water level judgment in train water tanks is solved, and the efficiency and accuracy of automated water supply is achieved, and water resources are avoided.
Patent Information
- Application Number
- CN202110845254.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-26
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-07-26
AI Technical Summary
The water level judgment of train water tanks is difficult to automate, which leads to the inability of water level to accurately judge, resulting in low water supply efficiency and easy waste of water resources.
Through the pre-established sound event detection model, the audio clips to be tested are obtained for feature extraction, and the deep learning model is used to predict the water level height order, and the audio signal is enhanced in combination with beamforming and noise reduction algorithms to achieve accurate positioning of the water level height.
Improve water supply efficiency, avoid waste of water resources caused by water overflow, and ensure the automation and accuracy of the water supply process.
Smart Images

Figure CN115691548B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of train intelligent control technology, and in particular to a method, device, equipment and storage medium for locating the water level of a train water tank. Background Art
[0002] During train travel, water tanks must be refilled periodically to meet the train's water needs. Automated train water-filling robots are a key development direction for replacing manual water-filling operations. However, since water-filling robots cannot automatically determine the water level in the tanks or the progress of water filling, automatic and efficient water-level monitoring is required to coordinate water-filling operations. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to propose a water level positioning method, device, equipment and storage medium for a train water tank, which can improve the water supply efficiency of the train water tank and avoid water overflow and waste of water resources.
[0004] Based on the above objectives, the first aspect of the present invention provides a method for locating the water level of a train water tank, comprising:
[0005] Get the audio clip to be tested;
[0006] Extracting features from the audio segment to be tested to obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested;
[0007] Substitute the multiple frames of feature vectors to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested.
[0008] Based on the same purpose, the second aspect of the present invention provides a train water tank water level positioning device, comprising:
[0009] An acquisition module is used to obtain the audio clip to be tested;
[0010] A feature extraction module is used to extract features from the audio segment to be tested, and obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested;
[0011] The water level magnitude obtaining module is used to substitute the multiple frames of feature vectors to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level magnitude corresponding to the audio segment to be tested.
[0012] Based on the same purpose, the third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0013] Based on the same purpose, the fourth aspect of the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any method described in the first aspect.
[0014] As can be seen from the above, the train water tank water level positioning method, device, equipment and storage medium provided by the present invention pre-establish a sound event detection model. After obtaining the audio segment to be tested, the audio segment to be tested is feature extracted to obtain multiple frames of feature vectors to be tested corresponding to the audio segment to be tested, and then the multiple frames of feature vectors to be tested are substituted into the sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested; by pre-establishing the sound event detection model, the correspondence between the feature vector corresponding to the audio segment and the water level height magnitude is pre-constructed. After obtaining the audio segment to be tested, its feature is first extracted to obtain the feature vector to be tested, and then the feature vector to be tested is substituted into the sound event detection model, which can quickly and accurately obtain the water level height magnitude corresponding to the feature vector to be tested, that is, obtain the water level height magnitude corresponding to the audio segment to be tested, which is convenient for timely judging the water level condition of the train water tank during the water filling process, improving the water filling efficiency and avoiding water overflow and waste of water resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 A schematic flow chart of a method for locating the water level of a train water tank provided by an embodiment of the present invention;
[0017] Figure 2 A schematic diagram of a flow chart of a method for establishing a sound event detection model provided by an embodiment of the present invention;
[0018] Figure 3 A schematic structural diagram of a train water tank water level positioning device provided by an embodiment of the present invention;
[0019] Figure 4 A more specific schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0021] It should be noted that, unless otherwise defined, technical or scientific terms used in the embodiments of the present invention should have the same ordinary meaning as those understood by persons of ordinary skill in the art to which the present invention belongs. Words such as "include" or "comprising" mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, but do not exclude other elements or objects.
[0022] During train travel, water tanks must be refilled periodically to meet the train's water needs. Automated train water-filling robots are a key development direction for replacing manual water-filling operations. However, since the water-filling robots themselves cannot automatically determine the water level in the tanks or the progress of water filling, to improve water filling efficiency while preventing overflow and water waste, an automatic and efficient water level monitoring system is required to assist the water-filling robots in water-filling operations.
[0023] To address the aforementioned issues, the present invention provides a method, device, equipment, and storage medium for locating the water level in a train water tank. A sound event detection model is pre-established. After obtaining a test audio clip, feature extraction is performed on the test audio clip to obtain a multi-frame feature vector corresponding to the test audio clip. The multi-frame feature vectors are then substituted into the sound event detection model to obtain the water level magnitude corresponding to the test audio clip. This method and device can be applied to mobile phones, tablets, computers, smart wearable devices, and the like, without limitation.
[0024] For ease of understanding, the train water tank water level positioning method is described in detail below with reference to the accompanying drawings.
[0025] Figure 1 A flow chart of a method for locating the water level of a train water tank provided by an embodiment of the present invention; Figure 1 As shown, the method includes the following steps:
[0026] S11. Obtain an audio clip to be tested.
[0027] In this step, the audio clip to be tested refers to an audio clip used to judge the water level in the water tank when the water-supply robot is filling water into the train water tank. The audio clip to be tested is the sound generated when the water-supply robot is filling water into the train water tank, which can be collected by the microphone carried by the water-supply robot. The water-supply robot can carry one microphone or multiple microphones, and there is no specific limitation; and a water-supply robot can be provided in each carriage of the train to fill water into the water tank of the carriage, so the audio clip to be tested can include only one audio clip or multiple audio clips, and there is no specific limitation.
[0028] In one case, it can be set that during the water filling process, when the water filling starts, the water filling robot will actively send an audio clip collected by the microphone it carries to the electronic device that executes this method (hereinafter referred to as this electronic device) as the audio clip to be tested for the electronic device to determine the water level in the water tank.
[0029] In one case, it can be set that during the water filling process, when the water filling starts, the electronic device sends an audio acquisition request to the water filling robot. Based on the received audio acquisition request, the water filling robot feeds back the audio segment collected by its microphone to the electronic device as the audio segment to be tested. After receiving the audio segment to be tested, the electronic device determines the water level in the water tank.
[0030] S12: Extract features from the audio segment to be tested, and obtain multiple frames of feature vectors to be tested corresponding to the audio segment to be tested.
[0031] Feature extraction refers to converting the original audio clip to be tested into a low-dimensional feature vector representation. In practical applications, Mel-frequency cepstral coefficients (MFCC) or linear prediction coefficients (LPCC) can be used to extract features from the audio clip to be tested, obtaining multiple frames of feature vectors corresponding to the audio clip to be tested. Taking the use of MFCC to extract features from the audio clip to be tested as an example, when MFCC is used to extract features from the audio clip to be tested, MFCC is used as the feature representation of each frame of audio. The audio clip to be tested can be represented as a multi-frame ordered MFCC feature vector, that is, the feature vector to be tested obtained after feature extraction is an MFCC feature vector.
[0032] In practical applications, after obtaining the audio segment to be tested, the audio segment to be tested may be preprocessed to remove noise from the audio segment to be tested and improve the purity of the audio segment to be tested. Therefore, in some possible implementations, before the step of extracting features from the audio segment to be tested and obtaining multiple frames of feature vectors to be tested corresponding to the audio segment to be tested, the step further includes: enhancing the audio segment to be tested.
[0033] Optionally, enhancing the audio segment to be tested includes: enhancing the audio segment to be tested by using a beamforming algorithm and a noise reduction algorithm.
[0034] Beamforming algorithms can enhance a microphone's sound pickup in a specific direction while suppressing noise from other directions. In practical applications, since the train's water tank is stationary, beamforming can enhance the audio captured by the water supply robot's microphone in the direction of the tank's location. Denoising algorithms can be used to suppress all other steady-state noise, excluding the sound of the train's water tank filling. Alternatively, the noise reduction algorithm can be first applied to the audio clip under test, followed by the beamforming algorithm. This is not a specific limitation.
[0035] It can be understood that by using beamforming algorithm and noise reduction algorithm to enhance the audio segment to be tested, the audio in the direction of the train water tank position can be enhanced while reducing steady-state noise. The result obtained when judging the water level of the train water tank based on the enhanced audio segment to be tested is more accurate.
[0036] S13. Substitute the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested.
[0037] In this step, the sound event detection model is a model pre-established using a deep learning model for predicting the water level of the train water tank based on the feature vector to be measured. It should be noted that the deep learning model refers to a sound event detection model based on deep learning.
[0038] In practical applications, in order to accurately locate the water level in the train water tank, a sound event detection model can be established in advance based on the deep learning model; Figure 2 A flow chart of a method for establishing a sound event detection model according to an embodiment of the present invention; Figure 2 As shown, in some possible implementations, the method for establishing a sound event detection model includes:
[0039] S21. Obtain training audio segments at different water level heights; wherein the training audio segments are manually intercepted from enhanced water level sound data at different water level heights;
[0040] S22, extracting features from the training audio clip to obtain a multi-frame training feature vector corresponding to the training audio clip;
[0041] S23, classifying the multiple frames of training feature vectors and obtaining the water level height magnitude corresponding to each frame of training feature vector;
[0042] S24. Use the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to train the deep learning model to obtain a sound event detection model.
[0043] Among them, the water level height magnitude refers to the height reached by the water level in the train water tank during the process of filling the train water tank with water using a water robot. The water level height magnitude can be set according to actual needs, and the water level height magnitude can be set to a discontinuous type. For example, the water level height magnitude can be set to "1 / 4 water level", "1 / 3 water level", "1 / 2 water level" or "1 water level", etc., and "1 water level" can be used to indicate that the water tank is full of water. There is no specific limitation.
[0044] The training audio clip refers to the audio clip used to train the deep learning model. It is the sound generated when the water supply robot fills the water tank of the train. It can also be collected by the microphone carried by the water supply robot. In actual applications, the water level sound data of the water supply of the train water tank collected by the microphone of the water supply robot may contain noise, and the audio stream of the water level sound data may contain other irrelevant sound waveforms. After obtaining the water level sound data of the water supply of the train water tank, the noise reduction algorithm and the beamforming algorithm can be used to enhance the water level sound data first. Then, the audio stream of the enhanced water level sound data can be manually intercepted to intercept all the waveforms containing the water level sound data and as few other irrelevant sound waveforms as possible at the beginning and end to obtain the training audio clip. After processing a water level sound data, multiple training audio clips can be obtained, or only one training audio clip can be obtained. There is no specific limitation.
[0045] For a given water level, the water level sound data can include the sound data of water tanks filling water in different train cars, or just the sound data of water tanks filling water in one car. When the water filling robot carries multiple microphones, the water level sound data can include the sound data of water tanks filling water collected by the robot's multiple microphones, or just the sound data of water tanks filling water collected by a single microphone of the robot, without specific limitations. Accordingly, for a given water level, the training audio clips can include audio clips of water tanks filling water in different train cars, or just the audio clip of water tanks filling water in one car, without specific limitations.
[0046] In one case, for the enhanced water level sound data, its audio stream can be displayed on the display interface of the electronic device, so as to facilitate manual interception of the audio stream through the display interface. After the interception is completed, a training audio clip is obtained, and then the confirmation button of the display interface can be clicked so that the electronic device can obtain the training audio clip and establish a sound event detection model based on the training audio clip.
[0047] After obtaining the training audio segment, feature extraction is first performed on the training audio segment. Then, in some possible implementations, feature extraction is performed on the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment, including: using Mel-frequency cepstral coefficients or linear prediction parameters to extract features from the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment.
[0048] For example, MFCC is used to extract features from a training audio segment, and MFCC is used as a feature representation of each frame of audio. The training audio segment can be represented as a multi-frame ordered MFCC feature vector.
[0049] Before training the deep learning model, the training feature vectors can be first divided into categories, and then the divided training feature vectors can be labeled with water level height magnitudes. The training feature vectors belonging to the same category are labeled with the same water level height magnitude, and the water level height magnitude corresponding to each frame of the training feature vector can be obtained as data for training the deep learning model.
[0050] For example, for a multi-frame training feature vector with a water level height magnitude of “1 / 2 water level”, all the training feature vectors may be classified into the same category and then labeled as “1 / 2 water level”.
[0051] In practical applications, when training a sound event detection model, for convenience, different numbers may be used to represent the corresponding water level height magnitudes. Therefore, in some possible implementations, before the step of training a deep learning model using multiple frames of training feature vectors corresponding to the training audio clips and the water level height magnitudes corresponding to each frame of the training feature vectors to obtain the sound event detection model, the step further includes:
[0052] The mapping relationship between the preset water level height magnitude and the marked number.
[0053] Among them, the marked numbers refer to numbers used to represent different water level height levels; the mapping relationship refers to the correspondence between the water level height level and the marked numbers, and the water level height level and the marked numbers are in a one-to-one correspondence.
[0054] In actual applications, the labeling numbers corresponding to different water level height levels can be set as needed; for example, the labeling number corresponding to the water level height level "0 water level" can be set to 0, the labeling number corresponding to the water level height level "1 / 8 water level" can be set to 1, the labeling number corresponding to the water level height level "1 / 6 water level" can be set to 2, the labeling number corresponding to the water level height level "1 / 2 water level" can be set to 3, or the labeling number corresponding to the water level height level "1 water level" can be set to 4, and so on, without specific limitation.
[0055] Correspondingly, the deep learning model is trained using multiple frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, including: the deep learning model is trained using multiple frame training feature vectors, the water level height magnitude corresponding to each frame training feature vector and the labeled numbers corresponding to each water level height magnitude to obtain a sound event detection model.
[0056] That is, based on the mapping relationship between water level height magnitude and labeled numbers, as well as the correspondence between training feature vectors and water level height magnitude, a deep learning model can be trained to obtain a sound event detection model, and a correspondence between training feature vectors and labeled numbers can be constructed.
[0057] It can be understood that the sound event detection technology and deep learning technology are combined to construct a sound event detection model. After obtaining the audio clip to be tested, its features are first extracted to obtain the feature vector to be tested, and then the feature vector to be tested is substituted into the sound event detection model to obtain the labeled number corresponding to the feature vector to be tested. Based on the labeled number, the water level height level corresponding to the feature vector to be tested can be determined, and then the water level height level corresponding to the audio clip to be tested can be determined, that is, the water level information corresponding to the audio clip to be tested is obtained.
[0058] In some possible implementations, substituting the feature vectors of multiple frames corresponding to the audio segment under test into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment under test includes:
[0059] Substitute the feature vector of each frame to be tested into the pre-established sound event detection model to obtain the label number corresponding to the feature vector of each frame to be tested;
[0060] According to the labeled numbers corresponding to the feature vectors to be measured in each frame, and based on the mapping relationship between the water level magnitude and the labeled numbers, the water level magnitude corresponding to the feature vectors to be measured in each frame is determined.
[0061] In practical applications, the deep learning model is a sound event detection model based on deep learning, comprising multiple serially connected time-delay neural networks and an artificial neural network classifier. When a characteristic vector to be tested is input into the sound event detection model, the serially connected time-delay neural networks can convert the multiple input characteristic vectors to be tested into corresponding multiple depth representation vectors, while the artificial neural network classifier can predict the labeled numbers corresponding to each frame of the characteristic vector to be tested based on the multiple depth representation vectors. Based on the correspondence between the water level height magnitude and the labeled numbers, the water level height magnitude corresponding to each frame of the characteristic vector to be tested is determined. When multiple characteristic vectors to be tested are input, a corresponding numerical sequence is determined, with each number in the numerical sequence representing a water level height magnitude.
[0062] It can be understood that by pre-establishing a sound event detection model and pre-constructing the correspondence between the feature vector corresponding to the audio clip and the water level height magnitude, after obtaining the audio clip to be tested, its features are first extracted to obtain the feature vector to be tested, and then the feature vector to be tested is substituted into the sound event detection model, which can quickly and accurately obtain the water level height magnitude corresponding to the feature vector to be tested, that is, the water level height magnitude corresponding to the audio clip to be tested, which is convenient for timely judgment of the water level situation of the train water tank during the water filling process, thereby improving the water filling efficiency and avoiding water overflow and waste of water resources.
[0063] In some possible implementations, the method further includes:
[0064] The water level height magnitude corresponding to the audio segment to be tested is sent to the control unit of the aquatic robot, so that the control unit of the aquatic robot determines whether to change the water supply pressure or stop water supply based on the water level height magnitude.
[0065] When the water level is low, the water robot can increase the water injection pressure and increase the water injection speed; when the water level is high, the water robot can reduce the water injection pressure and reduce the water injection speed; therefore, the control unit of the water robot can determine whether to stop filling water in time based on the received water level, thereby improving the water filling efficiency and avoiding water overflow and waste of water resources.
[0066] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the method.
[0067] It should be noted that the above description is limited to some embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0068] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides a water level positioning device for a train water tank.
[0069] refer to Figure 3 , the train water tank water level positioning device comprises:
[0070] An acquisition module 31 is used to acquire an audio segment to be tested;
[0071] A feature extraction module 32 is used to extract features from the audio segment to be tested and obtain a multi-frame feature vector to be tested corresponding to the audio segment to be tested;
[0072] The water level magnitude obtaining module 33 is used to substitute the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level magnitude corresponding to the audio segment to be tested.
[0073] In some possible implementations, the device further includes an enhancement module (not shown in the figure); the enhancement module is configured to enhance the audio segment to be tested by using a beamforming algorithm and a noise reduction algorithm.
[0074] In some possible implementations, the feature extraction module 32 is specifically configured to:
[0075] Mel-frequency cepstral coefficients or linear prediction parameters are used to extract features of the audio segment to be tested, and a multi-frame feature vector to be tested corresponding to the audio segment to be tested is obtained.
[0076] In some possible implementations, the device further includes a category labeling module (not shown in the figure) and a training module (not shown in the figure);
[0077] The category labeling module is used to label the category of multiple frames of training feature vectors and obtain the water level height magnitude corresponding to each frame of training feature vector;
[0078] A training module is used to train a deep learning model using multiple frames of training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame of training feature vectors to obtain a sound event detection model;
[0079] The acquisition module 31 is further configured to acquire training audio segments at different water level height levels; wherein the training audio segments are manually intercepted from the enhanced water level sound data at different water level height levels;
[0080] The feature extraction module 32 is further configured to extract features from the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment.
[0081] In some possible implementations, the feature extraction module 32 is further specifically configured to extract features from the training audio segment using Mel-frequency cepstral coefficients or linear prediction parameters to obtain a multi-frame training feature vector corresponding to the training audio segment.
[0082] In some possible implementations, the device further includes a mapping relationship building module (not shown in the figure);
[0083] Mapping relationship building module, specifically used for:
[0084] The mapping relationship between the preset water level magnitude and the marked numbers;
[0085] Training modules, specifically for:
[0086] The deep learning model is trained using multi-frame training feature vectors, the water level magnitude corresponding to each frame training feature vector, and the labeled numbers corresponding to each water level magnitude to obtain a sound time detection model.
[0087] In some possible implementations, the water level magnitude obtaining module 33 is specifically configured to:
[0088] Substitute the feature vector of each frame to be tested into the pre-established sound event detection model to obtain the label number corresponding to the feature vector of each frame to be tested;
[0089] According to the labeled numbers corresponding to the feature vectors to be measured in each frame, and based on the mapping relationship between the water level magnitude and the labeled numbers, the water level magnitude corresponding to the feature vectors to be measured in each frame is determined.
[0090] For ease of description, the above device is described separately based on its functionality within various modules. Of course, when implementing the present invention, the functionality of each module can be implemented within the same software or hardware components. The device of the above embodiment is used to implement the corresponding train water tank water level positioning method described in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be further elaborated here.
[0091] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the train water tank water level positioning method described in any of the above embodiments is implemented.
[0092] Figure 4 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0093] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0094] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0095] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0096] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0097] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0098] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0099] The electronic device of the above embodiment is used to implement the corresponding train water tank water level positioning method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0100] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the train water tank water level positioning method as described in any of the above embodiments.
[0101] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0102] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the train water tank water level positioning method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0103] It should be noted that the embodiments of the present invention may be further described in the following manner:
[0104] A method for locating the water level of a train water tank, comprising:
[0105] Get the audio clip to be tested;
[0106] Performing feature extraction on the audio segment to be tested to obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested;
[0107] Substitute the multiple frames of feature vectors to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested.
[0108] Optionally, before the step of extracting features from the audio segment to be tested and obtaining a plurality of frames of feature vectors to be tested corresponding to the audio segment to be tested, the method further includes:
[0109] enhancing the audio segment to be tested;
[0110] The step of enhancing the audio segment to be tested includes:
[0111] The audio segment to be tested is enhanced by using a beamforming algorithm and a noise reduction algorithm.
[0112] Optionally, the extracting features from the audio segment to be tested to obtain a plurality of frames of feature vectors to be tested corresponding to the audio segment to be tested includes:
[0113] Mel-frequency cepstral coefficients or linear prediction parameters are used to extract features of the audio segment to be tested, and a multi-frame feature vector to be tested corresponding to the audio segment to be tested is obtained.
[0114] Optionally, the method for establishing the sound event detection model includes:
[0115] Acquire training audio segments at different water level heights; wherein the training audio segments are manually intercepted from enhanced water level sound data at different water level heights;
[0116] Performing feature extraction on the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment;
[0117] The multi-frame training feature vectors are labeled and the water level height magnitude corresponding to each frame training feature vector is obtained;
[0118] The deep learning model is trained using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model.
[0119] Optionally, performing feature extraction on the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment includes:
[0120] Mel-frequency cepstral coefficients or linear prediction parameters are used to extract features from the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment.
[0121] Optionally, before the step of training a deep learning model using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, the method further includes:
[0122] The mapping relationship between the preset water level magnitude and the marked numbers;
[0123] The deep learning model is trained using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, including:
[0124] The deep learning model is trained using multi-frame training feature vectors, the water level magnitude corresponding to each frame training feature vector, and the labeled numbers corresponding to each water level magnitude to obtain a sound time detection model.
[0125] Optionally, substituting the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested includes:
[0126] Substitute the feature vector of each frame to be tested into the pre-established sound event detection model to obtain the label number corresponding to the feature vector of each frame to be tested;
[0127] According to the labeled numbers corresponding to the feature vectors to be measured in each frame, and based on the mapping relationship between the water level magnitude and the labeled numbers, the water level magnitude corresponding to the feature vectors to be measured in each frame is determined.
[0128] A water level positioning device for a train water tank, comprising:
[0129] An acquisition module is used to obtain the audio clip to be tested;
[0130] A feature extraction module is used to extract features from the audio segment to be tested, and obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested;
[0131] The water level magnitude obtaining module is used to substitute the multiple frames of feature vectors to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level magnitude corresponding to the audio segment to be tested.
[0132] An electronic device comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, a method for locating the water level of a train water tank is implemented.
[0133] A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to execute a method according to any one of claims 1 to 7.
[0134] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Within the scope of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.
[0135] In addition, to simplify the description and discussion, and in order not to obscure the embodiments of the present invention, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present invention, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present invention will be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that embodiments of the present invention may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0136] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0137] The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for locating the water level of a train water tank, characterized in that: include: Get the audio clip to be tested; Extracting features from the audio segment to be tested to obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested; Substituting the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested; The method for establishing the sound event detection model includes: Acquire training audio segments at different water level heights; wherein the training audio segments are manually intercepted from enhanced water level sound data at different water level heights; Performing feature extraction on the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment; The multi-frame training feature vectors are labeled and the water level height magnitude corresponding to each frame training feature vector is obtained; A deep learning model is trained using multiple frames of training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame of training feature vectors to obtain a sound event detection model; Before the step of training a deep learning model using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, the method further includes: The mapping relationship between the preset water level magnitude and the marked numbers; The deep learning model is trained using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, including: The deep learning model is trained using multi-frame training feature vectors, the water level magnitude corresponding to each frame training feature vector, and the labeled numbers corresponding to each water level magnitude to obtain a sound event detection model.
2. The train water tank water level positioning method according to claim 1, characterized in that: Before the step of extracting features from the audio segment to be tested and obtaining a plurality of frames of feature vectors to be tested corresponding to the audio segment to be tested, the method further includes: enhancing the audio segment to be tested; The step of enhancing the audio segment to be tested includes: The audio segment to be tested is enhanced by using a beamforming algorithm and a noise reduction algorithm.
3. The train water tank water level positioning method according to claim 1, characterized in that: The extracting features of the audio segment to be tested to obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested includes: Mel-frequency cepstral coefficients or linear prediction parameters are used to extract features of the audio segment to be tested, and a multi-frame feature vector to be tested corresponding to the audio segment to be tested is obtained.
4. The train water tank water level positioning method according to claim 1, characterized in that: Extracting features from the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment includes: Mel-frequency cepstral coefficients or linear prediction parameters are used to extract features from the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment.
5. The train water tank water level positioning method according to claim 1, characterized in that: Substituting the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested, including: Substitute the feature vector of each frame to be tested into the pre-established sound event detection model to obtain the label number corresponding to the feature vector of each frame to be tested; According to the labeled numbers corresponding to the feature vectors to be measured in each frame, and based on the mapping relationship between the water level magnitude and the labeled numbers, the water level magnitude corresponding to the feature vectors to be measured in each frame is determined.
6. A water level positioning device for a train water tank, characterized in that: include: An acquisition module is used to obtain the audio clip to be tested; A feature extraction module is used to extract features from the audio segment to be tested, and obtain a plurality of feature vectors to be tested corresponding to the audio segment to be tested; a water level magnitude obtaining module, configured to substitute the feature vectors of multiple frames to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level magnitude corresponding to the audio segment to be tested; The method for establishing the sound event detection model includes: Acquire training audio segments at different water level heights; wherein the training audio segments are manually intercepted from enhanced water level sound data at different water level heights; Performing feature extraction on the training audio segment to obtain a multi-frame training feature vector corresponding to the training audio segment; The multi-frame training feature vectors are labeled and the water level height magnitude corresponding to each frame training feature vector is obtained; A deep learning model is trained using multiple frames of training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame of training feature vectors to obtain a sound event detection model; Before the step of training a deep learning model using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, the method further includes: The mapping relationship between the preset water level magnitude and the marked numbers; The deep learning model is trained using the multi-frame training feature vectors corresponding to the training audio clips and the water level height magnitude corresponding to each frame training feature vector to obtain a sound event detection model, including: The deep learning model is trained using multi-frame training feature vectors, the water level magnitude corresponding to each frame training feature vector, and the labeled numbers corresponding to each water level magnitude to obtain a sound event detection model.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Equipment for detecting windshield washer fluid height in water tank by using sound, and windshield wiper system
CN109060077A
Tactile feedback method
CN109871120A