Voice text association method and device, electronic equipment and storage medium

By using a voice recognition and association system, the voice signal of the infrared thermal imager is converted into voice text and automatically associated with the thermal image data, which solves the problem that voice annotations cannot be automatically associated with thermal image data in the existing technology and improves analysis efficiency.

CN120853576APending Publication Date: 2025-10-28HANGZHOU MICROIMAGE SOFTWARE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511134172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing infrared thermal imagers cannot automatically link voice annotations with thermal image data, resulting in low efficiency in thermal image data analysis.

Method used

The speech signal is converted into speech text by a speech recognition engine, and the speech text is automatically associated with thermal image data by an association system. The location coordinates are obtained by combining the GPS module to generate a structured report.

Benefits of technology

It enables automatic association between voice text and thermal imaging data, improving the efficiency of thermal imaging data analysis and reducing the need for manual matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853576A_ABST
    Figure CN120853576A_ABST
Patent Text Reader

Abstract

The invention discloses a voice text association method and device, electronic equipment and a storage medium, belongs to the technical field of intelligent detection equipment, and is used for improving the association efficiency of a voice text and thermal image data. The method comprises the following steps: in response to a target instruction, converting an obtained language signal into a voice text; obtaining thermal image data to be associated with the voice text; and associating the voice text with the thermal image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intelligent detection equipment technology, specifically relating to a method, device, electronic device, and storage medium for associating voice and text. Background Technology

[0002] An infrared thermal imager is a device that uses infrared thermal imaging technology to detect the infrared radiation of a target object and, through signal processing and photoelectric conversion, converts the temperature distribution of the target object into a visual image. The infrared thermal imager accurately quantifies the actual detected heat and images the entire target object in real time in a planar format, forming thermal image data of the target object.

[0003] Some thermal imagers can record operators' voice notes during inspections, but they cannot correlate the recorded notes with the thermal image data of the target object. This requires manual matching of the voice notes and thermal image data during later analysis of the thermal image data, which greatly reduces the efficiency of thermal image data analysis. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for associating voice text with thermal imaging data, which can solve the problem that voice text used for annotation is not associated with thermal imaging data.

[0005] In a first aspect, embodiments of this application provide a method for associating speech and text, the method comprising: in response to a target instruction, converting an acquired language signal into speech and text; acquiring thermal image data to be associated with the speech and text; and associating the speech and text with the thermal image data.

[0006] Secondly, embodiments of this application provide a speech-text association device, which includes: a conversion module for converting acquired speech signals into speech-text in response to a trigger command; an acquisition module for acquiring thermal image data to be associated with the speech-text; and an association module for associating the speech-text and the thermal image data.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] In this embodiment of the application, by responding to the target instruction, the acquired language signal is converted into speech text; thermal image data to be associated with the speech text is acquired; and the speech text and the thermal image data are associated, the speech signal detected by the thermal imager can be converted into speech text during the inspection process, and then the thermal image data and the speech text can be associated, avoiding the need for manual matching of speech text and thermal image data in the later stage, and improving the association efficiency of thermal image data and speech text. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the system architecture of a thermal imager provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a method for associating voice and text provided in an embodiment of this application; Figure 3 This is a flowchart illustrating another method for associating voice and text provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a voice-text association device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0014] The following description, in conjunction with the accompanying drawings, details the speech-text association method, apparatus, electronic device, and storage medium provided in this application through specific embodiments and application scenarios.

[0015] One embodiment of this application provides a method for associating voice and text, which can be applied to thermal imagers. Figure 1 The system architecture of the thermal imager provided in the embodiments of this application is shown.

[0016] like Figure 1 As shown, the system architecture of the thermal imager mainly includes a hardware layer, a software layer, and a data layer. The hardware layer of the thermal imager includes an infrared camera with a resolution of 640×480 and a refresh rate of 25Hz. The thermal imager may also include a visible light camera with a video resolution of 1080P. Furthermore, the thermal imager may include a Global Positioning System (GPS) module with a positioning accuracy of ±1m.

[0017] like Figure 1 As shown, the hardware layer of the thermal imaging system architecture can also include a voice acquisition module, which can include a dual-microphone array with 120° directivity and a signal-to-noise ratio greater than or equal to 15dB. The dual-microphone array has strong noise suppression capabilities and is connected to a digital signal processor.

[0018] like Figure 1 As shown, the hardware layer of the thermal imaging system architecture can also include an edge computing unit capable of running speech models.

[0019] In this embodiment, the software layer of the thermal imager system architecture includes a speech recognition engine. This engine comprises an Automatic Speech Recognition (ASR) model that combines a Transformer architecture with offline deployment capabilities. The input to this model can be a speech signal, which is then converted into speech text and output. The ASR model includes industry terms such as descriptions of power equipment defects, like "bushing discharge" and "disconnector overheating." The software layer of the thermal imager system architecture may also include an NLP processing module. This NLP processing module, upon receiving the speech text output from the ASR model, matches thermal image data, such as the device number (e.g., B3-205A) and temperature value (e.g., 72°C), using regular expressions. The software layer of the thermal imager system architecture may also include an association system for associating speech text, thermal image data, and GPS location coordinates.

[0020] In this embodiment, the data layer of the thermal imager system architecture mainly includes thermal image data and voice text, and may also include structured reports generated based on thermal image data and voice text.

[0021] The thermal imaging system architecture in this embodiment features a dual-microphone array with strong noise suppression capabilities, which can improve the accuracy of speech recognition in industrial environments. The ASR speech model of the speech recognition engine includes industry terms such as power equipment defect description corpus, which can accurately identify the input speech and generate correct speech text, reducing the error rate of converting speech signals into speech text. Furthermore, the association system in this embodiment can associate speech text and thermal imaging data, avoiding the need for manual matching of thermal imaging data and speech text in the later stages, and improving the association efficiency between thermal imaging data and speech text.

[0022] Figure 2 This application illustrates a method for associating voice and text, as provided in an embodiment. Figure 2 As shown, the method includes the following steps: Step S201: In response to the target instruction, convert the acquired language signal into speech text.

[0023] Thermal imagers are primarily used in research and development, industrial inspection, and equipment maintenance, and also have wide applications in fire prevention, night vision, and security. Simply put, a thermal imager converts the invisible infrared energy emitted by an object into a visible thermal image. Different colors in the thermal image represent different temperatures of the object being measured. An infrared thermal imager is a device that uses infrared thermal imaging technology to detect the infrared radiation of a target object and, through signal processing and photoelectric conversion, converts the temperature distribution of the target object into a visible image. Infrared thermal imagers accurately quantify the actual detected heat and image the entire target object in real time in a planar format, thus accurately identifying suspected faulty areas that are overheating. Operators use the image colors and hotspot tracking display on the screen to initially judge the overheating situation and fault location, and then conduct rigorous analysis, demonstrating high efficiency and accuracy in confirming problems.

[0024] When using a thermal imager for inspection, if the imager detects an anomaly in a target object or needs to make notes on the target object, the user can speak the target command. The target object can be a building, electrical equipment, mechanical equipment, or other aspects. The target command can be a wake word, for example, the user can say the wake word "record defect". The thermal imager can detect the wake word spoken by the user. The detection threshold for the wake word by the thermal imager can be -20dB, and the false trigger rate is less than 0.1 times / hour.

[0025] After the thermal imager detects the target command (wake word) spoken by the user, it can activate the recording function. The recording function can have a sampling rate of 16kHz and a bit depth of 16bit. The recording function can collect the voice signal of the user's spoken notes.

[0026] In one embodiment, before converting the acquired speech signal into speech text, the method further includes: acquiring the speech signal through a dual-microphone array; filtering the speech signal through a power frequency noise filter; and improving the signal-to-noise ratio of the filtered speech signal through a preset beamforming algorithm.

[0027] In this embodiment, the recording function can acquire the voice signal of the user's spoken notes using a dual-microphone array. The dual microphones have strong noise suppression capabilities; by analyzing the time delay difference (TDOA) and amplitude difference between the two microphones, they can distinguish sounds from different directions. For example, when noise comes from the side, the array can suppress the signal from that direction while retaining the voice signal from directly in front. In this embodiment, the voice signal can also be filtered using a power frequency noise filter. Filtering the voice signal using a power frequency noise filter is a key step in voice signal processing to eliminate fixed-frequency interference (such as 50Hz / 60Hz noise). In this embodiment, suppressing noise in the voice signal using dual microphones and filtering the voice signal using a power frequency noise filter can suppress most of the noise in the voice signal.

[0028] In this embodiment, the signal-to-noise ratio of the filtered speech signal can also be improved by using a preset beamforming algorithm CNN-LSTM. By combining convolutional neural networks (CNN) and long short-term memory networks (LSTM), dynamic beamforming of the speech signal can be achieved. By learning spatial-temporal features to optimize the beam direction and suppress non-steady-state noise, the signal-to-noise ratio of the speech signal can be improved by 15dB.

[0029] In this embodiment, by using dual microphones, a power frequency noise filter, and a CNN-LSTM beamforming algorithm, noise in the speech signal can be effectively suppressed, thereby improving the accuracy of speech signal recognition in industrial environments.

[0030] In this embodiment of the application, VAD silence detection can be implemented: if there is continuous silence for more than or equal to 500ms when collecting user voice, the voice segment is cut off to complete the collection of user voice.

[0031] After acquiring the user's speech signal, it can be input into the ASR (Automatic Speech Recognition) model of the speech recognition engine. The ASR model converts the speech signal into speech-to-text. This ASR model is trained using a corpus of defect descriptions of target objects (electrical equipment, mechanical equipment, etc.), such as "bushing discharge" and "shunt switch overheating." This ASR model improves the accuracy of speech recognition, ensuring the accurate conversion of speech signals into speech-to-text and guaranteeing the accuracy of the resulting speech-to-text. Furthermore, the ASR model can enhance the differentiation of similar words through acoustic feature enhancement, such as distinguishing between "bushing" and "shunt switch," further improving the accuracy of speech recognition and ensuring the accurate conversion of speech signals into speech-to-text.

[0032] Step S202: Obtain thermal image data to be associated with the voice text.

[0033] In this embodiment, after converting the collected voice signal into voice-text, thermal image data to be associated with the voice-text can be obtained. Thermal image data refers to data collected by a thermal imager, including thermal images and temperature data. In this embodiment, the thermal image data to be associated with the voice-text is the thermal image data that the user wants to annotate. For example, if a user detects that the temperature of a device exceeds the standard using a thermal imager and needs to make a voice annotation about the device, then the thermal image data of that device is the thermal image data to be associated with the user's voice-text. Specifically, the thermal image data collected when the user's wake-up word is received can be used as the thermal image data to be associated with the voice-text, or the thermal image data to be associated with the voice-text can be obtained through timestamps (IEEE 1588v2 synchronization), that is, by matching the timestamps of the voice signal or voice-text with the timestamps of the thermal image data to determine the thermal image data to be associated with the voice-text.

[0034] In one embodiment, acquiring thermal image data to be associated with voice text includes: acquiring thermal image data based on a first timestamp of the voice signal, wherein the deviation between the second timestamp of the thermal image data and the first timestamp is within a time deviation range.

[0035] In this embodiment, thermal image data to be associated with speech text can be obtained via timestamps (IEEE 1588v2 synchronization). Specifically, upon receiving a speech signal, the first timestamp of that signal can be obtained. When associating speech text with thermal image data, if the second timestamp of the thermal image data and the first timestamp of the speech text are within the time deviation range, then the thermal image data is the one to be associated with the speech text. The time deviation range can be set according to actual needs, for example, ±50ms.

[0036] Step S203: Associate the voice text with the thermal image data.

[0037] In this embodiment, after obtaining the voice-to-text of the user's voice, a confidence level assessment can be performed on the voice-to-text. This assessment determines whether the voice-to-text recognition is accurate. If the confidence level meets the assessment criteria, vibration feedback can be provided and the voice-to-text can be saved. For example, if the confidence level (accuracy) of the voice-to-text is greater than 85%, it indicates that the confidence level meets the assessment criteria, and the voice-to-text can be saved. If the confidence level does not meet the assessment criteria, for example, if the confidence level (accuracy) is less than or equal to 85%, a reminder voice prompt, "Please describe again," can be played to remind the user to describe again.

[0038] If the confidence level of the voice text meets the judgment criteria, the voice text and thermal imaging data can be bound together to achieve the association between the voice text and thermal imaging data.

[0039] In one embodiment, after associating the voice text and the thermal image data, the method further includes: associating the voice text, the thermal image data, and the location coordinates at the time the thermal image data was acquired.

[0040] In other words, the embodiments of this application can also associate the GPS location coordinates detected by the GPS module when the thermal imager collects thermal image data with the voice text and thermal image data. In this way, the association between voice text, thermal image data and location coordinates can be realized, which makes it convenient for staff to quickly find the target object corresponding to the thermal image data and voice text based on the location coordinates.

[0041] As an example, here is a sample of binding and associating voice text and thermal imaging data: {Thermal image data: "110305.jpg", audio text: "Temperature abnormality in phase C of transformer No. 12", time: "2025-03-24T11:03:07+08:00", location coordinates: "28.6123N, 115.8934E"}.

[0042] As shown in the example above, the binding and association of voice text, thermal image data (thermal image, temperature) and location coordinates are realized. In this embodiment of the application, the voice text, thermal image data (thermal image, temperature) and location coordinates can also be associated with information such as the second timestamp (time) of the thermal image data.

[0043] The speech-text association method provided in this application converts the acquired language signal into speech-text in response to a target instruction; acquires thermal image data to be associated with the speech-text; and associates the speech-text and thermal image data. This method can convert the speech signal collected by the thermal imager into speech-text during the inspection process, and then associate the thermal image data detected by the thermal imager with the speech-text, avoiding the need for manual matching of speech-text and thermal image data in the later stage, and improving the association efficiency of thermal image data and speech-text.

[0044] In one embodiment, after associating the voice text and thermal image data with the location coordinates when the thermal image data was acquired, the method further includes: obtaining key data from the voice text and thermal image data; filling the key data and location coordinates into a structured template to generate a structured report.

[0045] In this embodiment, after associating voice text, thermal imaging data, and location coordinates, key data can be obtained from the voice text and thermal imaging data. This key data may include temperature, voice text content, target object number, etc., and can be set according to actual needs. After obtaining the key data, the key data and location coordinates can be filled into a structured template to generate a structured report. Users can then use this structured report to obtain the key data and analyze thermal imaging data, improving work efficiency.

[0046] In one embodiment, obtaining key data from the voice text and the thermal imaging data includes: obtaining the temperature of the target object and a second timestamp of the thermal imaging data from the thermal imaging data; obtaining the target object's ID and temperature anomalies of a target part of the target object from the voice text, wherein the key data includes the target object's temperature, the second timestamp, the target object's ID, and the temperature anomalies of the target part.

[0047] In this embodiment, key data may include the temperature of the target object, the second timestamp of the thermal imaging data, the target object's number, and the temperature anomaly of the target part of the target object. Specifically, the temperature of the target object and the second timestamp when the thermal imaging data was acquired can be obtained from the thermal imaging data. The target object's number and the temperature anomaly of the target part can be obtained from voice text. For example, if a user inputs voice text based on thermal imaging data: "The C phase of transformer No. 12 has an abnormally high temperature," the voice text can be used to obtain that the target object is a transformer, the target object's number is "No. 12," the target part of the target object is "C phase," and the temperature anomaly of the target part is "abnormally high temperature."

[0048] After obtaining key data such as the target object's temperature, the second timestamp of the thermal imaging data, the target object's number, and the temperature anomalies of the target part of the target object, the key data and location coordinates can be filled into a structured template to generate a structured report. In this way, users can obtain this key data based on the structured report, thereby analyzing the target object and thermal imaging data, which improves work efficiency.

[0049] In one embodiment, after obtaining key data from the voice text and the thermal image data, the method further includes: if the obtained key data is missing target key data, obtaining the target key data input by the user.

[0050] In this embodiment, after acquiring key data from voice text or thermal imaging data, if it is found that the target key data is missing, the user can manually select or input the missing target key data to complete the key data. Specifically, when the acquired key data is missing target key data, the thermal imager can pop up a quick fill window, where the user can manually select or input the missing target key data. For example, if the missing target key data is the target object's number, a fill drop-down menu can pop up in the quick fill window, allowing the user to manually select the target object's number. Another example is if the actual target key data is the target object's temperature, which the user can input via the temperature keypad in the quick fill window. This allows for handling of missing key data anomalies, ensuring the integrity of the key data.

[0051] Below, through Figure 3 The specific examples shown illustrate the detailed process of the voice-text association method provided in this application embodiment: When a user inspects a product using a handheld thermal imager, the user says the wake word "record defect". The thermal imager then performs a wake word detection (threshold: -20dB, false trigger rate ≤0.1 times / hour). If the wake word is detected, the thermal imager's recording function is activated (sampling rate: 16KHz, bit depth 16bit).

[0052] This application embodiment can implement VAD silence detection: when collecting user voice, if continuous silence lasts for more than or equal to 500ms, the voice segment is cut to complete the user voice acquisition. After acquiring the user's voice signal, the voice signal can be input into the ASR model of the speech recognition engine. In the ASR model, the voice signal can be converted into speech-to-text, and the actual response of the ASR model is less than 200ms. After obtaining the speech-to-text of the user's voice, the confidence level of the speech-to-text can be judged. If the confidence level of the speech-to-text meets the judgment condition, vibration feedback is provided and the speech-to-text is saved. If the confidence level of the speech-to-text does not meet the judgment condition, a reminder voice "Please describe again" can be played to remind the user to describe again. When the confidence level of the speech-to-text meets the judgment condition, thermal image data to be associated with the speech-to-text can be acquired, and the speech-to-text, thermal image data, and position coordinates when the thermal image data was acquired can be bound together to realize the association between speech-to-text, thermal image data, and position coordinates.

[0053] After binding and associating voice text, thermal imaging data, and location coordinates, key data can be extracted from the voice text and thermal imaging data. If the key data is incomplete, a pop-up window will appear, allowing the user to input the key data. If the key data is complete, a structured template can be populated based on the key data and location coordinates to output a structured report.

[0054] It should be noted that the speech-text association method provided in this application embodiment can be executed by a speech-text association device or a control module within that device for executing the speech-text association method. This application embodiment uses the execution of the speech-text association method by a speech-text association device as an example to illustrate the speech-text association device provided in this application embodiment.

[0055] Figure 4 This is a schematic diagram of the structure of a voice-text association device according to an embodiment of this application. For example... Figure 4 As shown, the voice-text association device 400 includes: a conversion module 410, an acquisition module 420, and an association module 430.

[0056] The conversion module 410 is used to convert the acquired voice signal into voice text in response to a trigger command; the acquisition module 420 is used to acquire thermal image data to be associated with the voice text; and the association module 430 is used to associate the voice text and the thermal image data.

[0057] In one embodiment, the conversion module 410 is further configured to acquire the speech signal through a dual-microphone array; filter the speech signal through a power frequency noise filter; and improve the signal-to-noise ratio of the filtered speech signal through a preset beamforming algorithm.

[0058] In one embodiment, the acquisition module 420 is used to acquire the thermal image data based on the first timestamp of the voice signal, wherein the deviation between the second timestamp of the thermal image data and the first timestamp is within the time deviation range.

[0059] In one embodiment, the association module 430 is further configured to associate the voice text, the thermal image data, and the location coordinates when the thermal image data was acquired.

[0060] In one embodiment, the association module 430 is further configured to obtain key data from the voice text and the thermal image data; fill the key data and the location coordinates into a structured template to generate a structured report.

[0061] In one embodiment, the association module 430 is used to obtain the temperature of the target object and the second timestamp of the thermal imaging data from the thermal imaging data; and to obtain the number of the target object and the temperature anomaly of the target part of the target object from the voice text. The key data includes the temperature of the target object, the second timestamp, the number of the target object, and the temperature anomaly of the target part.

[0062] In one embodiment, the association module 430 is further configured to acquire the target key data input by the user when the acquired key data is missing target key data.

[0063] The voice-text association device in this application embodiment can be a device, or a component or integrated circuit in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. A non-mobile electronic device can be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not impose specific limitations.

[0064] The voice-text association device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0065] The voice-text association device provided in this application embodiment can achieve Figures 1 to 3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0066] like Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they perform the following: in response to a target instruction, converting the acquired language signal into speech text; acquiring thermal image data to be associated with the speech text; and associating the speech text and the thermal image data.

[0067] In one embodiment, before the voice signal obtained by the conversion is voice text, the voice signal is acquired through a dual-microphone array; the voice signal is filtered by a power frequency noise filter; and the signal-to-noise ratio of the filtered voice signal is improved by a preset beamforming algorithm.

[0068] In one embodiment, the thermal image data is obtained based on a first timestamp of the voice signal, wherein the deviation between the second timestamp of the thermal image data and the first timestamp is within a time deviation range.

[0069] In one embodiment, after associating the voice text and the thermal image data, the voice text, the thermal image data, and the location coordinates at the time the thermal image data was acquired are associated.

[0070] In one embodiment, after associating the voice text and the thermal image data with the location coordinates when the thermal image data was acquired, key data is obtained from the voice text and the thermal image data; the key data and the location coordinates are then filled into a structured template to generate a structured report.

[0071] In one embodiment, the temperature of the target object and a second timestamp of the thermal imaging data are obtained from the thermal imaging data; the number of the target object and the temperature anomaly of the target part of the target object are obtained from the voice text, and the key data includes the temperature of the target object, the second timestamp, the number of the target object, and the temperature anomaly of the target part.

[0072] In one embodiment, after obtaining key data from the voice text and the thermal image data, if the obtained key data is missing target key data, the target key data input by the user is obtained.

[0073] The specific execution steps can be found in the various steps of the above-described embodiment of the voice-text association method, and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0074] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.

[0075] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.

[0076] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).

[0077] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.

[0078] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described speech-text association method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0079] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk.

[0080] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0081] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0082] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for associating speech and text, characterized in that, include: In response to the target instruction, the acquired language signal is converted into speech text; Acquire thermal image data to be associated with the spoken text; The voice text and the thermal image data are associated.

2. The association method according to claim 1, characterized in that, Before the converted speech signal becomes speech text, the process further includes: The voice signal is acquired using a dual-microphone array; The speech signal is filtered using a power frequency noise filter; The signal-to-noise ratio of the filtered speech signal is improved by using a preset beamforming algorithm.

3. The method according to claim 1, characterized in that, The step of acquiring thermal image data to be associated with the speech text includes: The thermal image data is obtained based on the first timestamp of the voice signal, wherein the deviation between the second timestamp of the thermal image data and the first timestamp is within the time deviation range.

4. The association method according to claim 1, characterized in that, Following the association of the speech text and the thermal image data, the method further includes: The voice text, the thermal image data, and the location coordinates when the thermal image data was acquired are associated.

5. The association method according to claim 4, characterized in that, After associating the voice text and the thermal image data with the location coordinates when the thermal image data was acquired, the method further includes: Key data are obtained from the spoken text and the thermal image data; Fill the key data and location coordinates into the structured template to generate a structured report.

6. The association method according to claim 5, characterized in that, The step of obtaining key data from the speech text and the thermal image data includes: The temperature of the target object and a second timestamp of the thermal image data are obtained from the thermal image data; The target object's ID and the temperature anomaly of the target object's target location are obtained from the voice text. The key data includes the target object's temperature, the second timestamp, the target object's ID, and the temperature anomaly of the target location.

7. The association method according to claim 5, characterized in that, After obtaining key data from the speech text and the thermal image data, the method further includes: If the acquired key data is missing the target key data, the target key data input by the user is acquired.

8. A device for associating voice and text, characterized in that, include: The conversion module is used to convert the acquired voice signal into voice-to-text in response to a trigger command. The acquisition module is used to acquire thermal image data to be associated with the voice text; The association module is used to associate the voice text and the thermal image data.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the speech-text association method as described in any one of claims 1-5.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the speech-text association method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method and apparatus for managing images using voice tag

    CN105512164A

  • Method and device for generating infrared thermal image analysis report and manufacturing template

    CN115146603A

  • Method and system for generating structured report based on voice data, and storage medium

    CN118609747A

  • Report generation method and device, electronic equipment, storage medium and program product

    CN119849463A

  • A infrared thermal imager for infrared temperature measurement

    CN208383314U