Subtitle matching and display method and system based on audio file processing

The target personnel area is determined through infrared sensors, the demand coefficient is calculated based on the historical startup time of the Internet of Things device, and the audio files and subtitles are generated, which solves the problem that the audio content and prompt operations cannot be explained in detail in the prior art, and realizes the convenient operation of the Internet of Things device.

CN119922372BActive Publication Date: 2025-07-18WUXI FUTURE MIRROR DISPLAY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510092497.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-07-18
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The prior art cannot explain the audio content in more detail and serve as a reminder to viewers, and cannot effectively prompt hearing-impaired people and ordinary users to correctly operate IoT devices.

Method used

The area where the target person is located is determined through infrared sensors, the demand coefficient is calculated based on the historical startup time of the Internet of Things device, and audio files and subtitles are generated to prompt the operation.

Benefits of technology

It improves the convenience of using IoT devices, and by comprehensively considering the startup probability and urgency, it generates accurate subtitle prompts, improving the accuracy and training efficiency of the text generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922372B_ABST
    Figure CN119922372B_ABST
Patent Text Reader

Abstract

The present invention provides a subtitle matching and display method and system based on audio file processing, which relates to the technical field of audio processing. The method includes: determining a target area where a target person is located, determining a demand coefficient of an Internet of Things device according to a historical startup time of the Internet of Things device in the target area, and further determining a target Internet of Things device; determining an audio file according to the type and function of the target Internet of Things device; determining a function menu according to the audio file and the type and function of the target Internet of Things device; generating recommended setting information according to the audio file, the type and function of the target Internet of Things device, and selection buttons in the function menu; generating recommended text information according to the recommended setting information and the audio file, and further generating subtitles. According to the present invention, subtitles can be generated when the Internet of Things device plays an audio file, so as to give operation prompts to the target person, facilitate the target person to correctly operate the Internet of Things device, and improve the usability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly relates to a subtitle matching display method and system based on audio file processing. Background Art

[0002] In the related art, CN106504754B discloses a method for generating real-time subtitles according to audio output, and the steps are as follows: for the audio information that needs to be output by an electronic device, the following operations are performed: an audio acquisition module is used to monitor the audio information output by the electronic device in real time and collect it; the collected audio information is transmitted to a speech extraction module, and irrelevant content such as background music in the audio information is filtered and noise reduction processing is performed to obtain accurate speech information; then the speech information that needs to be converted into text is input into a speech recognition module to obtain the text information corresponding to the speech; finally, the converted text is displayed on the device screen in real time in the form of subtitles through a display module. The advantages of this solution are: it can help hearing-impaired people obtain the speech content contained in videos, audios or other forms, provides an effective and convenient way for hearing-impaired people to obtain speech information, and also provides convenience for ordinary users.

[0003] CN117809654B discloses a method, device, electronic device and medium for generating low-resource audio subtitles. By applying this technical solution, in a multi-modal pre-training model including a language encoder and an audio encoder, first, a language decoder is trained for the existing language encoder by using text data with a relatively sufficient sample size. And in the subsequent process, the language encoder is replaced with an audio encoder to indirectly train a language decoder for the audio encoder. So that a high-precision audio multi-modal pre-training model can be trained only with audio pairing data with a small sample size. Thus, a technical solution is realized that can still achieve high model performance even when there is only a small amount of available training audio-subtitle data pairs.

[0004] Therefore, using the related technology, the content in the audio can be converted into text to generate subtitles, but the related technology can only make the subtitles consistent with the content of the audio, and cannot explain the content of the audio in more detail, nor can it play a reminder role for viewers.

[0005] The information disclosed in the background art part of the present application is only intended to deepen the understanding of the general background art of the present application, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0006] The present invention provides a subtitle matching display method and system based on audio file processing, which can solve the technical problem that the related technology cannot explain the content of the audio in more detail.

[0007] According to a first aspect of the present invention, there is provided a subtitle matching display method based on audio file processing, including:

[0008] Determine a target area where a target person is located among a plurality of areas through an infrared sensor provided on an Internet of Things device, wherein each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component;

[0009] Determine a demand coefficient of the target person for a plurality of Internet of Things devices according to the historical startup time of the Internet of Things devices in the target area;

[0010] Determine a target Internet of Things device for starting the screen and the audio component according to the demand coefficient;

[0011] Determine an audio file to be played through the audio component according to the type and function of the target Internet of Things device;

[0012] Determine a function menu to be displayed on the screen according to the audio file, and the type and function of the target Internet of Things device;

[0013] Generate recommended setting information according to the audio file, the type and function of the target Internet of Things device, and selection buttons in the function menu;

[0014] Generate recommended text information according to the recommended setting information and the audio file;

[0015] Determine subtitles to be displayed in the function menu according to the recommended text information.

[0016] According to a second aspect of the present invention, there is provided a subtitle matching display system based on audio file processing, including:

[0017] A target area module, configured to determine a target area where a target person is located among a plurality of areas through an infrared sensor provided on an Internet of Things device, wherein each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component;

[0018] A demand coefficient module, configured to determine a demand coefficient of the target person for a plurality of Internet of Things devices according to the historical startup time of the Internet of Things devices in the target area;

[0019] A target Internet of Things device module, configured to determine a target Internet of Things device for starting the screen and the audio component according to the demand coefficient;

[0020] An audio file module, configured to determine an audio file to be played through the audio component according to the type and function of the target Internet of Things device;

[0021] A function menu module, configured to determine a function menu displayed on a screen according to the audio file, as well as the type and functions of the target Internet of Things device;

[0022] A recommended setting information module, configured to generate recommended setting information according to the audio file, the type and functions of the target Internet of Things device, and selection buttons in the function menu;

[0023] A recommended text information module, configured to generate recommended text information according to the recommended setting information and the audio file;

[0024] A subtitle module, configured to determine subtitles displayed in the function menu according to the recommended text information.

[0025] By adopting the above technical solutions, the present invention can achieve the following technical effects:

[0026] According to the present invention, the demand coefficient of a target person for multiple Internet of Things devices can be determined based on the historical startup time of the Internet of Things device by the target person, so as to determine the Internet of Things device to be started. When the Internet of Things device plays an audio file to prompt the target person to operate, subtitles can be generated to prompt the target person to operate, so as to facilitate the target person to correctly operate the Internet of Things device and improve the usability. When determining the demand coefficient, the demand degree of the target person for the Internet of Things device can be comprehensively described from two aspects: the probability of starting the Internet of Things device after the target person enters the target area and the eagerness to start the Internet of Things device, so as to obtain the demand coefficient of the Internet of Things device and improve the comprehensiveness and accuracy of the demand coefficient. When determining the startup demand coefficient, the overall demand of the current environment for starting various functions can be described by the maximum value of the average relative deviation and the average relative change of the environmental parameters corresponding to various functions, so as to improve the objectivity and accuracy of the startup demand coefficient. When training the text generation model, a conditional function can be used to determine whether there is an error in the overall semantics of the sample recommended text, and the capabilities of different aspects of the text generation model can be improved respectively in the case of no error and the case of error, that is, the text polishing ability of the text generation model and the description accuracy of the text are improved, so as to improve the accuracy of the text generation model and improve the training efficiency.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present invention. According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present invention will be clearer. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings;

[0029] Figure 1 Exemplarily shown is a schematic flowchart of a subtitle matching display method based on audio file processing according to an embodiment of the present invention;

[0030] Figure 2 Exemplarily shown is a block diagram of a subtitle matching display system based on audio file processing according to an embodiment of the present invention. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0032] The following will specifically describe the technical solutions of the present invention with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0033] Figure 1 Exemplarily shown is a schematic flowchart of a subtitle matching display method based on audio file processing according to an embodiment of the present invention, and the method includes:

[0034] Step S101, determine the target area where the target person is located among multiple areas through an infrared sensor set on the Internet of Things device, where each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component;

[0035] Step S102, determine the demand coefficient of the target person for multiple Internet of Things devices according to the historical startup time of the Internet of Things devices in the target area;

[0036] Step S103, determine the target Internet of Things device for starting the screen and the audio component according to the demand coefficient;

[0037] Step S104, determine the audio file to be played through the audio component according to the type and function of the target Internet of Things device;

[0038] Step S105: Determine the function menu displayed on the screen according to the audio file, as well as the type and functions of the target Internet of Things device.

[0039] Step S106: Generate recommended setting information according to the audio file, the type and functions of the target Internet of Things device, and the selection buttons in the function menu.

[0040] Step S107: Generate recommended text information according to the recommended setting information and the audio file.

[0041] Step S108: Determine the subtitles to be displayed in the function menu according to the recommended text information.

[0042] According to the subtitle matching and display method based on audio file processing in an embodiment of the present invention, the demand coefficient of a target person for multiple Internet of Things devices can be determined through the historical startup time of the Internet of Things devices by the target person, so as to determine the started Internet of Things device. When the Internet of Things device plays an audio file to prompt the target person to operate, subtitles can be generated to prompt the target person to operate, so as to facilitate the target person to correctly operate the Internet of Things device and improve the usability.

[0043] According to an embodiment of the present invention, in step S101, the Internet of Things device may include devices such as household appliances, for example, air conditioners, speakers, humidifiers, etc. The area may be a room. Each area may include at least one Internet of Things device. Among one or more Internet of Things devices in the same area, at least one Internet of Things device includes an infrared sensor, which can detect whether the target person is in this area. If the target person is in this area, this area is determined as the target area. And the Internet of Things device may include a screen and an audio component. The screen can display a specific picture and can also receive the touch input of the target person. The audio component may include a sound playback component (such as a speaker) for playing an audio file, and a sound receiving component (such as a microphone) for receiving a sound signal. For example, the sound receiving component can receive the voice input instruction of the target person and can also be used to detect the noise in the environment, etc.

[0044] According to an embodiment of the present invention, in step S102, a demand coefficient can be used to represent the usage demand of a target person for Internet of Things devices in a target area. Determining the demand coefficients of the target person for multiple Internet of Things devices according to the historical startup times of the Internet of Things devices in the target area includes: determining multiple target time periods when the target person is in the target area within multiple historical cycles; determining the historical startup times of each Internet of Things device within each target time period; determining the startup order of each Internet of Things device within each target time period according to the historical startup times of each Internet of Things device within each target time period; determining whether each Internet of Things device is started within each target time period; and determining the demand coefficients of each Internet of Things device according to the startup order of each Internet of Things device within each target time period and whether each Internet of Things device is started within each target time period.

[0045] According to an embodiment of the present invention, each historical cycle can be one day, one week, etc., and the present invention does not limit the duration of the historical cycle. Multiple target time periods when the target person is in the target area can be determined. For example, within the past multiple days, multiple target time periods when the target person is in the room can be determined.

[0046] According to an embodiment of the present invention, within each target time period, the historical startup times of each Internet of Things device can be determined, and based on the order of the historical startup times, the startup order of the Internet of Things devices within each target time period can be determined. For example, it can be determined which Internet of Things devices are started each time the target person enters the room and the startup order of these Internet of Things devices.

[0047] According to an embodiment of the present invention, determining the demand coefficients of each Internet of Things device according to the startup order of each Internet of Things device within each target time period and whether each Internet of Things device is started within each target time period includes: determining the demand coefficient D of the i-th Internet of Things device according to formula (1) i ,

[0048]

[0049] where N is the number of target time periods, N start,i is the number of times the i-th Internet of Things device is started within multiple target time periods, M is the number of Internet of Things devices in the target area, and O i,j is the startup order of the i-th Internet of Things device within the j-th target time period, i ≤ M, j ≤ N, and i, M, j, and N are all positive integers.

[0050] According to an embodiment of the present invention, in formula (1), For the start-up times of the i-th Internet of Things device and the number of target time periods, it can be used as the ratio of the number of times the i-th Internet of Things device is started after the target person enters the target area to the total number of times the target person enters the target area. The higher this ratio, the higher the probability that the target person starts the i-th Internet of Things device after entering the target area, and it can also indicate that the target person has the highest demand for the i-th Internet of Things device.

[0051] According to an embodiment of the present invention, in formula (1), It can be used to represent the eagerness of the target person to start the i-th Internet of Things device after entering the target area for the j-th time. That is, the earlier the start order (the smaller the value of O i,j ), the larger this ratio. Therefore, this ratio can be used to describe the eagerness of the target person to start the i-th Internet of Things device through the start order of the i-th Internet of Things device, and it can also indicate the demand degree of the target person for the i-th Internet of Things device. Then it can represent the average value of the eagerness of the target person to start the i-th Internet of Things device after entering the target area multiple times, which can be used to describe the overall eagerness of the target person to start the i-th Internet of Things device, and it can also describe the demand degree of the target person for the i-th Internet of Things device.

[0052] According to an embodiment of the present invention, multiplying the probability that the target person starts the i-th Internet of Things device after entering the target area by the overall eagerness of the target person to start the i-th Internet of Things device can obtain the demand coefficient of the i-th Internet of Things device, which can be used to describe the demand degree of the target person for the i-th Internet of Things device. Based on the above method, the demand coefficient of each Internet of Things device can be determined.

[0053] In this way, the demand degree of the target person for the Internet of Things device can be comprehensively described from two aspects: the probability that the target person starts the Internet of Things device after entering the target area and the eagerness to start the Internet of Things device, so as to obtain the demand coefficient of the Internet of Things device and improve the comprehensiveness and accuracy of the demand coefficient.

[0054] According to an embodiment of the present invention, in step S103, the target Internet of Things device for starting the screen and audio components can be determined based on the demand coefficient. For example, each Internet of Things device can be sorted according to the demand coefficient of each Internet of Things device, and the Internet of Things device ranked first can be used as the target Internet of Things device. After the Internet of Things device ranked first is set up, the Internet of Things device ranked second can be used as the target Internet of Things device, and so on.

[0055] According to an embodiment of the present invention, in step S104, the screen and audio components of the target Internet of Things device can be started, and based on the type and function of the target Internet of Things device, the audio file played by the audio component is determined. For example, the audio file is a prompt audio message for prompting the target person to set the target Internet of Things device.

[0056] According to an embodiment of the present invention, determining the audio file played through the audio component according to the type and function of the target Internet of Things device includes: starting the corresponding environmental parameter detection component according to the type of the target Internet of Things device; at multiple moments within a preset time period, obtaining at least one type of environmental parameter through the environmental parameter detection component; determining the start demand coefficient of the function of the type corresponding to the environmental parameter according to at least one type of environmental parameter; determining the function to be started according to the start demand coefficient of each function; generating a prompt message according to the function to be started, and generating an audio file corresponding to the prompt message; playing the audio file through the audio component.

[0057] According to an embodiment of the present invention, the corresponding environmental parameter detection component can be started according to the type of the target Internet of Things device. The environmental parameter detection component can be set on the target Internet of Things device, or the environmental parameter detection components corresponding to multiple Internet of Things devices can be centrally set together. The present invention does not limit this. In an example, if the type of the target Internet of Things device is an air conditioner, the environmental parameter detection component may include an air temperature detection component and an air humidity detection component. If the type of the target Internet of Things device is a speaker, the environmental parameter detection component is a noise detection component. The present invention does not limit the specific type of the environmental parameter detection component.

[0058] According to an embodiment of the present invention, the preset time period can be 10 seconds, 30 seconds, 1 minute, etc. The present invention does not limit the duration of the preset time period. The time interval between adjacent moments is 1 second, 0.1 second, etc. The present invention does not limit the time interval between adjacent moments. The environmental parameter can be obtained through the environmental parameter detection component. As in the above example, the air temperature and air humidity at multiple moments can be obtained, or the noise intensity at multiple moments can be obtained.

[0059] According to an embodiment of the present invention, the start demand coefficient can be used to describe whether the environment in the target area is suitable. If the environment is not suitable, for example, the air humidity is high, or the air temperature is high, then the demand for starting the dehumidification function of the air conditioner for dehumidification, or starting the temperature adjustment function of the air conditioner for cooling is high.

[0060] According to an embodiment of the present invention, determining a startup demand coefficient for a function of a type corresponding to an environmental parameter according to at least one type of environmental parameter includes: determining the startup demand coefficient D for the k-th type of function corresponding to the environmental parameter according to formula (2). start,k ,

[0061]

[0062] where p k,s,t is the measured value of the s-th environmental parameter corresponding to the k-th type of function at the t-th moment, p p,k,s is the preset value of the s-th environmental parameter corresponding to the k-th type of function, p k,s,t+1 is the measured value of the s-th environmental parameter corresponding to the k-th type of function at the (t + 1)-th moment, max is the maximum value function, m is the number of types of environmental parameters corresponding to the k-th type of function, n is the number of moments within a preset time period, and s, t, m, and n are all positive integers.

[0063] According to an embodiment of the present invention, in formula (2), is the relative deviation between the s-th environmental parameter corresponding to the k-th type of function (for example, the air temperature corresponding to the temperature adjustment function) and the preset value, is the average relative deviation of multiple environmental parameters corresponding to the k-th type of function at multiple moments. The larger this average relative deviation, the more the environmental parameters corresponding to the k-th type of function deviate from the appropriate situation, and the higher the demand for starting the k-th function.

[0064] According to an embodiment of the present invention, in formula (2), represents the relative change of environmental parameters at adjacent moments, represents the average relative change of multiple environmental parameters at multiple moments. The larger this average relative change, the worse the overall stability of the multiple environmental parameters. For example, the noise rapidly increases, or the noise fluctuates greatly, etc., and the higher the demand for starting the k-th function.

[0065] According to an embodiment of the present invention, the maximum value of the above average relative deviation and average relative change can be taken as the startup demand coefficient for the k-th type of function to describe the overall demand for the k-th type of function.

[0066] In this way, the overall demand of the current environment for starting various functions can be described by the maximum value of the average relative deviation and average relative change of the environmental parameters corresponding to various functions, improving the objectivity and accuracy of the startup demand coefficient.

[0067] According to an embodiment of the present invention, the function with the highest startup requirement coefficient can be determined as the function to be started. After the function with the highest startup requirement coefficient is set, the function with the second highest startup requirement coefficient can be determined as the function to be started, and so on. Of course, multiple functions can also be set at one time. For example, one or more functions with a startup requirement coefficient greater than or equal to a preset coefficient threshold can be determined as the functions to be started. For example, the temperature adjustment function and the dehumidification function of the air conditioner can both be set as the functions to be started.

[0068] According to an embodiment of the present invention, a prompt message can be generated according to the function to be started, and an audio file corresponding to the prompt message can be generated. For example, in the case of a relatively high air temperature, a prompt message can be generated by combining the detected air temperature and the name of the function to be started. For example, a text-type prompt message such as "It is detected that the temperature is high, and it is recommended to start the temperature adjustment function to cool down" can be generated through a text generation model, and an audio file can be generated using this prompt message, that is, an audio file that reads out the pronunciation of each word in the prompt message, and the audio file can be played through an audio component.

[0069] According to an embodiment of the present invention, in step S105, according to the audio file, as well as the type and function of the target Internet of Things device, the function menu to be displayed on the screen is determined. That is, the prompt message corresponding to the audio file may involve one or more functions, and the function menu to be displayed can be selected from the multiple functions of the target Internet of Things device based on these functions. For example, among the function menus for setting the temperature of the air conditioner, the function menu for setting the dehumidification function, and the function menu for setting the ventilation function, the function menu for setting the temperature and the function menu for setting the dehumidification function to be displayed can be selected.

[0070] According to an embodiment of the present invention, in step S106, based on the prompt message corresponding to the audio file, the type of function to be recommended for setting can be determined, that is, the name of the target function to be started. The setting method can also be determined based on the selection button in the function menu, such as entering the number of the temperature, or the button for raising or lowering the temperature, etc. The recommended value of the environmental parameter can also be determined, such as the preset value of the above environmental parameter. That is, recommended setting information including information such as the target function, environmental parameter, setting method, and recommended value is obtained. For example, the recommended setting information for recommending to start the temperature adjustment function of the air conditioner and adjusting the current air temperature value of 30°C to the recommended value of 25°C by inputting temperature data.

[0071] According to an embodiment of the present invention, in step S107, the recommended text information is the text that can be displayed on the screen to prompt the target person to make settings. By using a specific method to display the recommended text information, subtitles can be obtained.

[0072] According to an embodiment of the present invention, based on the recommended setting information and the audio file, recommended text information is generated, including: determining, according to the recommended setting information, a target function to be started of a target Internet of Things device, environmental parameters corresponding to the target function, and recommended values of the environmental parameters set for recommendation; processing the target function, the environmental parameters corresponding to the target function, and the recommended values of the environmental parameters through a trained text generation model to obtain a to-be-determined recommended text; obtaining the audio text of the audio file; verifying the to-be-determined recommended text through the audio text; if the verification passes, determining the to-be-determined recommended text as the recommended text information; otherwise, retraining the text generation model.

[0073] According to an embodiment of the present invention, from the above-mentioned recommended setting information, the target function to be started, the environmental parameters corresponding to the target function, and the recommended values of the environmental parameters set for recommendation can be parsed, and can be processed through a trained text generation model to obtain a to-be-determined recommended text, that is, a text obtained by fusing and polishing the name of the target function, the current environmental parameters, and the recommended values of the environmental parameters. For example, "It is detected that the current temperature is relatively high, reaching 30°C, and the comfort level is not good. It is recommended that you use the temperature adjustment function and set the temperature to 25°C to reduce the temperature and improve the comfort level."

[0074] According to an embodiment of the present invention, the text generation model can be a deep learning neural network model. For example, a recurrent neural network model. The present invention does not limit the specific type of the text generation model. Before performing the above-mentioned processing using the text generation model, training can be performed.

[0075] According to an embodiment of the present invention, the training steps of the text generation model include: obtaining a sample function, sample environmental parameters corresponding to the sample function, and sample recommended values; processing the sample function, the sample environmental parameters, and the sample recommended values through the text generation model to obtain a sample recommended text; determining a reference adjustment direction description text according to the sample environmental parameters and the sample recommended values; determining a reference environmental description text according to the sample environmental parameters; determining a reference environmental adjustment description text according to the sample recommended values; determining a sample adjustment direction description text, a sample environmental description text, and a sample environmental adjustment description text according to the sample recommended text; obtaining a verification text according to the sample function, the sample environmental parameters corresponding to the sample function, and the sample recommended values; determining a loss function of the text generation model according to the verification text, the sample recommended text, the reference adjustment direction description text, the reference environmental description text, the reference environmental adjustment description text, the sample adjustment direction description text, the sample environmental description text, and the sample environmental adjustment description text; training the text generation model according to the loss function of the text generation model to obtain a trained text generation model.

[0076] According to an embodiment of the present invention, the sample function can be any function of any Internet of Things device, the sample environmental parameter can be a measured environmental parameter or an artificially set environmental parameter, the sample recommended value is the recommended value of the artificially set environmental parameter, and the above information can be processed by a text generation model to obtain a sample recommended text.

[0077] According to an embodiment of the present invention, the above information and the sample recommended text can be analyzed to determine the accuracy of the sample recommended text and thus determine the accuracy of the text generation model. For example, based on the sample environmental parameter and the sample recommended value, a reference adjustment direction description text can be determined. For example, if the sample environmental parameter is 30°C and the sample recommended value is 25°C, the reference adjustment direction description text is "decrease", "lower the temperature", etc. Also, based on the sample environmental parameter, a reference environment description text can be determined. For example, if the sample environmental parameter is 30°C, the reference environment description text is a text describing a relatively high environmental temperature, such as "high temperature", "poor comfort", "high air temperature", etc. Further, based on the sample recommended value, a reference environment adjustment description text can be determined. For example, if the sample recommended value is 25°C, the reference environment adjustment description text is "cool", "lower the temperature", "improve comfort", etc.

[0078] According to an embodiment of the present invention, from the sample recommended text, a sample adjustment direction description text, a sample environment description text, and a sample environment adjustment description text can be determined, that is, texts describing the adjustment direction, the current environment, and the environment after adjustment are parsed from the sample recommended text.

[0079] According to an embodiment of the present invention, based on the sample function, the sample environmental parameter corresponding to the sample function, and the sample recommended value, a verification text can also be obtained. The verification text can be connected with these three pieces of information in simple statements without polishing, for example, "The current temperature is 30°C. It is recommended that you use the temperature adjustment function and set the temperature to 25°C".

[0080] According to an embodiment of the present invention, based on the above information, the error in the sample recommended text generated by the text generation model can be determined, and thus the loss function of the text generation model can be determined through this error.

[0081] According to an embodiment of the present invention, determining a loss function of a text generation model according to the verification text, the sample recommendation text, the reference adjustment direction description text, the reference environment description text, the reference environment adjustment description text, the sample adjustment direction description text, the sample environment description text, and the sample environment adjustment description text includes: determining, through a semantic recognition model, the verification semantic information of the verification text, the sample recommendation semantic information of the sample recommendation text, the reference adjustment direction semantic information of the reference adjustment direction description text, the reference environment semantic information of the reference environment description text, the reference environment adjustment semantic information of the reference environment adjustment description text, the sample adjustment direction semantic information of the sample adjustment direction description text, the sample environment semantic information of the sample environment description text, and the sample environment adjustment semantic information of the sample environment adjustment description text; determining the loss function LOSS of the text generation model according to formula (3),

[0082]

[0083] where S RD is the reference adjustment direction semantic information, S SD is the sample adjustment direction semantic information, (S RD ) T is the transposed vector of S RD , sim p is a preset similarity threshold, S RE is the reference environment semantic information, (S RE ) T is the transposed vector of S RE , S SE is the sample environment semantic information, S RA is the reference environment adjustment semantic information, (S RA ) T is the transposed vector of S RA , S SA is the sample environment adjustment description text, S SR is the sample recommendation semantic information, S V is the verification semantic information, w1 and w2 are preset weights, and if is a conditional function.

[0084] According to an embodiment of the present invention, the semantic recognition model may be a recurrent neural network model, which can be used to determine the semantic information in vector form of various information, that is, the sample recommendation semantic information of the sample recommendation text, the reference adjustment direction semantic information of the reference adjustment direction description text, the reference environment semantic information of the reference environment description text, the reference environment adjustment semantic information of the reference environment adjustment description text, the sample adjustment direction semantic information of the sample adjustment direction description text, the sample environment semantic information of the sample environment description text, and the sample environment adjustment semantic information of the sample environment adjustment description text.

[0085] According to an embodiment of the present invention, in formula (3), the conditional function indicates that when the conditional function value is w1|S RE -S SE | + w2|S RA -S SA |; otherwise, the conditional function value is It means that the cosine similarity between the reference adjustment direction semantic information and the sample adjustment direction semantic information is higher than or equal to a preset similarity threshold, that is, the semantic similarity between the reference adjustment direction description text and the sample adjustment direction description text is relatively high, and the two are synonyms. That is, the sample adjustment direction description text in the sample recommendation text generated by the text generation model is correct. For example, the meanings of the sample adjustment direction description text and the reference adjustment direction description text are both to lower the temperature. In this case, the overall semantics of the sample recommendation text is not incorrect, and only the text describing the current environmental parameters and the adjusted environmental parameters (for example, the text for polishing) needs to be optimized, that is, to improve the text polishing ability of the text generation model. The errors between the reference environmental semantic information and the sample environmental semantic information, and the errors between the reference environmental adjustment semantic information and the sample environmental adjustment description text can be weighted and summed, that is, the semantic errors of the text describing the current environmental parameters and the adjusted environmental parameters are weighted and summed to obtain the conditional function value as the loss function, and the loss function is reduced during the training process, so that the semantic errors of the text describing the current environmental parameters and the adjusted environmental parameters are reduced, and the text polishing ability of the text generation model is improved.

[0086] According to an embodiment of the present invention, if the cosine similarity between the reference adjustment direction semantic information and the sample adjustment direction semantic information is lower than the preset similarity threshold, that is, the semantic similarity between the reference adjustment direction description text and the sample adjustment direction description text is relatively low, that is, the sample adjustment direction description text in the sample recommendation text generated by the text generation model is incorrect. For example, the meaning of the reference adjustment direction description text is to lower the temperature, and the meaning of the sample adjustment direction description text is to raise the temperature. In this case, the overall semantics of the sample recommendation text is incorrect. It can be through |S SR -S VDescribe the overall semantic error of the sample recommendation text, and use the cosine similarity between the reference environment semantic information and the sample environment semantic information, the cosine similarity between the reference environment adjustment semantic information and the sample environment adjustment description text, and the cosine similarity between the reference adjustment direction semantic information and the sample adjustment direction semantic information as the denominator to amplify the overall semantic error. Take the amplified semantic error as the loss function, which can enhance the training intensity. Moreover, during the training process, the loss function can be reduced, so that the semantic error between the sample recommendation semantic information and the verification semantic information is reduced, improving the semantic accuracy of the sample recommendation text. It can also improve the cosine similarity between the reference environment semantic information and the sample environment semantic information, the cosine similarity between the reference environment adjustment semantic information and the sample environment adjustment description text, and the cosine similarity between the reference adjustment direction semantic information and the sample adjustment direction semantic information, thereby improving the description accuracy of the text generation model for the adjustment direction and the description accuracy for the current environmental parameters and the adjusted environmental parameters.

[0087] According to an embodiment of the present invention, the above loss function can be backpropagated, and the parameters of the text generation model can be adjusted by the gradient descent method. The above adjustment steps can be executed multiple times to perform multiple trainings to obtain a trained text generation model.

[0088] In this way, the conditional function can be used to determine whether there is an error in the overall semantics of the sample recommendation text. In the cases of no error and error respectively, different aspects of the text generation model can be improved, that is, improving the text polishing ability of the text generation model and the description accuracy of the text, thereby improving the accuracy of the text generation model and enhancing the training efficiency.

[0089] According to an embodiment of the present invention, using the trained text generation model, a pending recommendation text can be generated. Further, the pending recommendation text can be verified, that is, the audio text of the audio text can be used for verification. The audio text is the text composed of multiple words in the audio file, and its meaning is theoretically similar to the pending recommendation text, but the description of specific data is not exact. That is, the audio text is the text corresponding to the audio text for prompting the target person to set, and only has a prompting effect. For example, it prompts the target person that the current temperature is high and the temperature adjustment function can be turned on for cooling, but does not involve more details, such as the current specific environmental parameters and the recommended values of the environmental parameters.

[0090] According to an embodiment of the present invention, verifying the to-be-recommended text through an audio text includes: determining a target function description text, a current environment description text, and an environment adjustment description text according to the to-be-recommended text information; forming a comparison text by the target function description text, the current environment description text, and the environment adjustment description text; determining comparison semantic information of the comparison text through a semantic recognition model; determining audio semantic information of the audio text through the semantic recognition model; determining the semantic similarity between the comparison semantic information and the audio semantic information; and determining that the verification is passed when the semantic similarity is greater than or equal to a preset semantic similarity threshold.

[0091] For example, the target function description text, the current environment description text, and the environment adjustment description text can be formed into a comparison text. For example, "The current temperature is high, the comfort level is poor. It is recommended that you use the temperature adjustment function to lower the temperature and improve the comfort level", and its comparison semantic information can be obtained. Further, the audio semantic information of the audio text "It is detected that the temperature is high. It is recommended to start the temperature adjustment function to cool down" can also be determined. The comparison semantic information and the audio semantic information are theoretically highly similar. For example, it is greater than or equal to the preset semantic similarity threshold. If the semantic similarity is greater than or equal to the preset semantic similarity threshold, it is determined that the verification is passed, and the to-be-recommended text is determined as the recommended text information. On the contrary, the text generation model can be retrained, and the performance of the text generation model can be further improved by using richer training data, so as to generate a more accurate to-be-recommended text.

[0092] According to an embodiment of the present invention, in step S108, the caption to be displayed in the function menu can be determined according to the recommended text information. For example, the font, font size, color, etc. of the recommended text information can be set to obtain the caption, and it is displayed in the function menu on the screen.

[0093] The subtitle matching and display method based on audio file processing according to an embodiment of the present invention can determine the demand coefficients of a target person for multiple Internet of Things devices based on the historical startup times of the Internet of Things devices by the target person, so as to determine the Internet of Things devices to be started. When the Internet of Things devices play audio files to prompt the target person to operate, subtitles can be generated to prompt the target person to operate, so that the target person can operate the Internet of Things devices correctly and improve the usability. When determining the demand coefficients, the demand degree of the target person for the Internet of Things device can be comprehensively described from two aspects: the probability of starting the Internet of Things device after the target person enters the target area and the urgency of starting the Internet of Things device, so as to obtain the demand coefficients of the Internet of Things devices and improve the comprehensiveness and accuracy of the demand coefficients. When determining the startup demand coefficients, the overall demand of the current environment for starting various functions can be described by the maximum value of the average relative deviation and the average relative change of the environmental parameters corresponding to various functions, so as to improve the objectivity and accuracy of the startup demand coefficients. When training the text generation model, a conditional function can be used to determine whether there is an error in the overall semantics of the sample recommended text, and the capabilities of different aspects of the text generation model can be improved respectively when there is no error and when there is an error, that is, the text polishing ability of the text generation model and the description accuracy of the text can be improved, so as to improve the accuracy of the text generation model and the training efficiency.

[0094] Figure 2 Exemplarily shown is a block diagram of a subtitle matching and display system based on audio file processing according to an embodiment of the present invention. The system includes:

[0095] A target area module, configured to determine the target area where the target person is located in multiple areas through an infrared sensor disposed on the Internet of Things device. Each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component;

[0096] A demand coefficient module, configured to determine the demand coefficients of the target person for multiple Internet of Things devices according to the historical startup times of the Internet of Things devices in the target area;

[0097] A target Internet of Things device module, configured to determine the target Internet of Things device for starting the screen and the audio component according to the demand coefficients;

[0098] An audio file module, configured to determine the audio file played through the audio component according to the type and function of the target Internet of Things device;

[0099] A function menu module, configured to determine the function menu displayed on the screen according to the audio file and the type and function of the target Internet of Things device;

[0100] A recommended setting information module, configured to generate recommended setting information according to the audio file, the type and function of the target Internet of Things device, and the selection buttons in the function menu;

[0101] A recommended text information module, configured to generate recommended text information according to the recommended setting information and the audio file;

[0102] A subtitle module, configured to determine the subtitles to be displayed in the function menu according to the recommended text information.

[0103] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The object of the present invention has been fully and effectively achieved. The function and structural principle of the present invention have been demonstrated and explained in the embodiments. Without departing from the principle, any deformation or modification can be made to the embodiments of the present invention.

[0104] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A subtitle matching and display method based on audio file processing, characterized in that, Including: An infrared sensor disposed on an Internet of Things device is used to determine a target area where a target person is located among multiple areas. Each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component; According to the historical start times of the Internet of Things devices in the target area, determine the demand coefficients of the target person for multiple Internet of Things devices; According to the demand coefficients, determine the target Internet of Things devices for starting the screen and audio components; According to the type and function of the target Internet of Things device, determine the audio file to be played through the audio component; According to the audio file, and the type and function of the target Internet of Things device, determine the function menu displayed on the screen; According to the audio file, the type and function of the target Internet of Things device, and the selection buttons in the function menu, generate recommended setting information; According to the recommended setting information and the audio file, generate recommended text information; According to the recommended text information, determine the subtitles to be displayed in the function menu; Determining the audio file to be played through the audio component according to the type and function of the target Internet of Things device includes: According to the type of the target Internet of Things device, start the corresponding environmental parameter detection component; At multiple moments within a preset time period, obtain at least one type of environmental parameter through the environmental parameter detection component; According to at least one type of environmental parameter, determine the start demand coefficient of the function corresponding to the environmental parameter type; According to the start demand coefficients of each function, determine the function to be started; According to the function to be started, generate a prompt message and generate an audio file corresponding to the prompt message; Play the audio file through the audio component.

2. The subtitle matching and display method based on audio file processing according to claim 1, wherein Determining the demand coefficients of the target person for multiple Internet of Things devices according to the historical start times of the Internet of Things devices in the target area includes: Within multiple historical cycles, determine multiple target time periods when the target person is in the target area; Determine the historical start times of each Internet of Things device within each target time period; According to the historical start times of each Internet of Things device within each target time period, determine the start order of each Internet of Things device within each target time period; Determine whether each Internet of Things device is started within each target time period; According to the start order of each Internet of Things device within each target time period, and whether each Internet of Things device is started within each target time period, determine the demand coefficients of each Internet of Things device.

3. The subtitle matching and display method based on audio file processing according to claim 2, characterized in that Determining the demand coefficients of each Internet of Things device according to the start order of each Internet of Things device within each target time period, and whether each Internet of Things device is started within each target time period includes: According to the formula Determine the demand coefficient D of the i-th Internet of Things device i , where N is the number of target time periods, N start,i is the number of startups of the i-th Internet of Things device in multiple target time periods, M is the number of Internet of Things devices in the target area, O i,j is the startup order of the i-th Internet of Things device in the j-th target time period, i ≤ M, j ≤ N, and i, M, j, and N are all positive integers.

4. The subtitle matching and display method based on audio file processing according to claim 1, wherein Determining the start demand coefficient of the function corresponding to the environmental parameter type according to at least one type of environmental parameter includes: According to the formula Determine the activation requirement coefficient D for the k-th type of function corresponding to the environmental parameter start,k , where p k,s,t is the measured value of the s-th environmental parameter corresponding to the k-th type of function at the t-th moment, p p,k,s is the preset value of the s-th environmental parameter corresponding to the k-th type of function, p k,s,t+1 is the measured value of the s-th environmental parameter corresponding to the k-th type of function at the (t + 1)-th moment, max is the maximum value function, m is the number of types of environmental parameters corresponding to the k-th type of function, n is the number of moments within the preset time period, and s, t, m, and n are all positive integers.

5. The subtitle matching and display method based on audio file processing according to claim 1, wherein Generating recommended text information according to the recommended setting information and the audio file includes: According to the recommended setting information, determine the target function to be started of the target Internet of Things device, the environmental parameter corresponding to the target function, and the recommended value of the environmental parameter of the recommended setting; Process the target function, the environmental parameters corresponding to the target function, and the recommended values of the environmental parameters through a trained text generation model to obtain a pending recommended text; Obtain the audio text of the audio file; Verify the pending recommended text through the audio text; If the verification passes, determine the pending recommended text as the recommended text information; Otherwise, retrain the text generation model.

6. The subtitle matching and display method based on audio file processing according to claim 5, characterized in that The training steps of the text generation model include: Obtain a sample function, sample environmental parameters corresponding to the sample function, and sample recommended values; Process the sample function, the sample environmental parameters, and the sample recommended values through the text generation model to obtain a sample recommended text; Determine a reference adjustment direction description text according to the sample environmental parameters and the sample recommended values; Determine a reference environmental description text according to the sample environmental parameters; Determine a reference environmental adjustment description text according to the sample recommended values; Determine a sample adjustment direction description text, a sample environmental description text, and a sample environmental adjustment description text according to the sample recommended text; Obtain a verification text according to the sample function, the sample environmental parameters corresponding to the sample function, and the sample recommended values; Determine the loss function of the text generation model according to the verification text, the sample recommended text, the reference adjustment direction description text, the reference environmental description text, the reference environmental adjustment description text, the sample adjustment direction description text, the sample environmental description text, and the sample environmental adjustment description text; Train the text generation model according to the loss function of the text generation model to obtain a trained text generation model.

7. The subtitle matching and display method based on audio file processing according to claim 6, characterized in that Determine the loss function of the text generation model according to the verification text, the sample recommended text, the reference adjustment direction description text, the reference environmental description text, the reference environmental adjustment description text, the sample adjustment direction description text, the sample environmental description text, and the sample environmental adjustment description text, including: Determine the verification semantic information of the verification text, the sample recommended semantic information of the sample recommended text, the reference adjustment direction semantic information of the reference adjustment direction description text, the reference environmental semantic information of the reference environmental description text, the reference environmental adjustment semantic information of the reference environmental adjustment description text, the sample adjustment direction semantic information of the sample adjustment direction description text, the sample environmental semantic information of the sample environmental description text, and the sample environmental adjustment semantic information of the sample environmental adjustment description text through a semantic recognition model; According to the formula Determine the loss function LOSS of the text generation model, where S RD is the semantic information of the reference adjustment direction, S SD is the semantic information of the sample adjustment direction, (S RD ) T is the transposed vector of S RD , sim p is the preset similarity threshold, S RE is the semantic information of the reference environment, (S RE ) T is the transposed vector of S RE , S SE is the semantic information of the sample environment, S RA is the semantic information of the reference environment adjustment, (S RA ) T is the transposed vector of S RA , S SA is the description text of the sample environment adjustment, S SR is the semantic information of the sample recommendation, S V is the verification semantic information, w1 and w2 are preset weights, and if is a conditional function; Conditional function Indicates that when the case is, the conditional function value is w1|S RE -S SE | + w2|S RA -S SA |, otherwise, the conditional function value is 8. The subtitle matching and display method based on audio file processing according to claim 5, characterized in that Verify the pending recommended text through the audio text, including: Determine a target function description text, a current environmental description text, and an environmental adjustment description text according to the pending recommended text information; Form a comparison text with the target function description text, the current environmental description text, and the environmental adjustment description text; Determine the comparison semantic information of the comparison text through a semantic recognition model; Determine the audio semantic information of the audio text through a semantic recognition model; Determine the semantic similarity between the comparison semantic information and the audio semantic information; Determine that the verification passes when the semantic similarity is greater than or equal to a preset semantic similarity threshold.

9. A subtitle matching and display system based on audio file processing, characterized in that, Including: A target area module, configured to determine a target area where a target person is located among multiple areas through an infrared sensor disposed on an Internet of Things device, where each area includes at least one Internet of Things device, and the Internet of Things device includes a screen and an audio component; A demand coefficient module, configured to determine a demand coefficient of the target person for multiple Internet of Things devices according to the historical start time of the Internet of Things devices in the target area; A target Internet of Things device module, configured to determine a target Internet of Things device for starting the screen and the audio component according to the demand coefficient; An audio file module, configured to determine an audio file to be played through the audio component according to the type and function of the target Internet of Things device; A function menu module, configured to determine a function menu displayed on the screen according to the audio file, and the type and function of the target Internet of Things device; A recommended setting information module, configured to generate recommended setting information according to the audio file, the type and function of the target Internet of Things device, and the selection buttons in the function menu; A recommended text information module, configured to generate recommended text information according to the recommended setting information and the audio file; A subtitle module, configured to determine subtitles to be displayed in the function menu according to the recommended text information; Determining an audio file to be played through the audio component according to the type and function of the target Internet of Things device includes: Starting a corresponding environment parameter detection component according to the type of the target Internet of Things device; Obtaining at least one type of environment parameter through the environment parameter detection component at multiple moments within a preset time period; Determining a start demand coefficient of a function corresponding to the environment parameter according to at least one type of environment parameter; Determining a function to be started according to the start demand coefficients of each function; Generating a prompt message according to the function to be started, and generating an audio file corresponding to the prompt message; Playing the audio file through the audio component.

Citation Information

Patent Citations

  • A method for real-time subtitle generation based on audio output

    CN106504754B

  • Low-resource audio subtitle generation methods, devices, electronic equipment and media

    CN117809654B

  • Household appliance control method and device, electronic equipment and storage medium

    CN116627045A

  • Device and method for providing audiovisual content for disabled person

    WO2023101159A1