Speech transcription method, speech transcription apparatus, and electronic device

By displaying the speech waveform and text information on the screen respectively, and adjusting the decibel range to filter noise based on user input, the problem of quickly identifying noise and blank segments in speech recordings is solved, simplifying operation and improving noise filtering efficiency.

CN114822546BActive Publication Date: 2026-02-27VIVO MOBILE COMM CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210350766.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2026-02-27
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

In existing technologies, noise and blank audio segments are difficult to quickly identify during the transcription of voice recordings, and noise reduction operations are complex and difficult for users to operate.

Method used

The system displays the waveform of the speech and the transcribed text information in two areas of the screen, respectively, and receives user input to determine the decibel range, filter noise, and update the display.

Benefits of technology

It simplifies user operation, improves the visualization analysis of voice information and noise filtering efficiency, reduces the difficulty of operation, and achieves intuitive noise filtering effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114822546B_ABST
    Figure CN114822546B_ABST
Patent Text Reader

Abstract

The application discloses a speech transcription method, a speech transcription device and an electronic device, and belongs to the technical field of communication. The speech transcription method comprises the following steps: based on a first speech, displaying a waveform graph corresponding to the first speech in a first area of a display screen, and displaying text information obtained by transcribing the first speech in a second area of the display screen; receiving a first input of a user; in response to the first input, determining a decibel range; filtering the first speech according to the decibel range, and updating the display of the waveform graph and the text information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of communication, and particularly relates to a voice transcription method, a voice transcription device and an electronic device. BACKGROUND

[0002] With the maturity of voice transcription technology, recording voice information in different scenes by using a voice transcription device has been widely applied.

[0003] In daily life, people will use a recording device to record relevant voice information in different scenes such as interviews, live broadcasts and speeches. However, due to the complexity of voice input scenes, a large amount of noise and blank audio segments will appear in the process of voice transcription of the recording material, and it is difficult to quickly determine effective information from the recording material. In the related technology, professional software is usually used to perform noise reduction processing on a computer, and the operation is complex and difficult to master. SUMMARY

[0004] The purpose of the embodiments of the application is to provide a voice transcription method, a voice transcription device and an electronic device, which can solve the problem of complex noise reduction operation.

[0005] In a first aspect, the embodiments of the application provide a voice transcription method, which comprises:

[0006] displaying, based on a first voice, a waveform graph corresponding to the first voice in a first area of a display screen and displaying text information obtained by transcribing the first voice in a second area of the display screen;

[0007] receiving a first input of a user;

[0008] determining a decibel range in response to the first input;

[0009] filtering the first voice according to the decibel range and updating the display of the waveform graph and the text information.

[0010] In a second aspect, the embodiments of the application provide a voice transcription device, which comprises:

[0011] a display module configured to display, based on a first voice, a waveform graph corresponding to the first voice in a first area of a display screen and display text information obtained by transcribing the first voice in a second area of the display screen;

[0012] a first receiving module configured to receive a first input of a user;

[0013] a first processing module configured to determine a decibel range in response to the first input;

[0014] filter the first voice according to the decibel range and update the display of the waveform graph and the text information.

[0015] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. The programs or instructions, when executed by the processor, implement the steps of the method according to the first aspect.

[0016] In a fourth aspect, a readable storage medium is provided, which stores programs or instructions. The programs or instructions, when executed by a processor, implement the steps of the method according to the first aspect.

[0017] In a fifth aspect, a chip is provided, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute programs or instructions to implement the method according to the first aspect.

[0018] In a sixth aspect, a computer program product is provided, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the method according to the first aspect.

[0019] In the embodiments of the present application, the waveform graph of the first voice and the transcribed text information are displayed on two interfaces respectively, and the noise of the waveform graph and the text information is filtered by adjusting the decibel range to update the first voice. The display is more intuitive, and the operation difficulty of the user is greatly reduced. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is a flowchart of the voice transcription method provided by the embodiments of the present application;

[0021] Figure 2 is one of the interface diagrams of the voice transcription method provided by the embodiments of the present application;

[0022] Figure 3 is another of the interface diagrams of the voice transcription method provided by the embodiments of the present application;

[0023] Figure 4 is a third of the interface diagrams of the voice transcription method provided by the embodiments of the present application;

[0024] Figure 5 is a fourth of the interface diagrams of the voice transcription method provided by the embodiments of the present application;

[0025] Figure 6 is a structural diagram of the voice transcription device provided by the embodiments of the present application;

[0026] Figure 7 is a structural diagram of the electronic device provided by the embodiments of the present application;

[0027] Figure 8 FIG. 1 is a hardware schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0029] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents a "or" relationship between the front and rear associated objects.

[0030] The speech transcription method and the speech transcription device provided by the embodiments of the present application will be described in detail below with reference to the drawings and specific embodiments and their application scenarios.

[0031] The speech transcription method can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.

[0032] The terminal includes, but is not limited to, a mobile phone or a tablet computer and other portable communication devices having a touch-sensitive surface (for example, a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the terminal can not be a portable communication device, but a desktop computer having a touch-sensitive surface (for example, a touch screen display and / or a touchpad).

[0033] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. It should be understood that the terminal can include one or more other physical user interface devices such as physical keyboards, mice, and joysticks.

[0034] The speech transcription method provided by the embodiments of the present application, the execution subject of the speech transcription method can be an electronic device or a function module or function entity capable of realizing the speech conversion method in the electronic device, the electronic device mentioned in the embodiments of the present application includes but is not limited to mobile phones, tablet computers, notebook computers, wearable devices, etc., the speech conversion method provided by the embodiments of the present application will be described below taking the electronic device as an execution subject.

[0035] As shown in the figure, the voice transcription method comprises steps 110, 120 and 130. Figure 1

[0036] Step 110, based on the first voice, displaying a waveform graph corresponding to the first voice in a first area of the display screen, and displaying text information obtained by transcribing the first voice in a second area of the display screen.

[0037] It should be noted that the first voice refers to the voice of the text information to be transcribed. The voice can be a recording material of offline scenes such as meetings and speeches, for example, recording the speeches of multiple participants in a meeting with a recording device, and then transcribing the recorded voice into text information after the meeting, and then sorting out; it can also be the sound information of online activities such as network live broadcast or online learning classroom, for example, recording the entire sound content of an online class with a voice transcription device, and viewing in the form of text to avoid missing some sound segments due to distraction or other reasons during online class; it can also be other sounds that need to be recorded and transcribed or transcribed in real time, and the present embodiment does not make specific limitations.

[0038] In this step, the first area and the second area of the display screen refer to two independent display areas on the display screen interface, which are used to display the waveform graph and the text information corresponding to the first voice.

[0039] It should be noted that the waveform graph corresponding to the first voice refers to a characteristic analysis graph obtained by converting the voice signal through voice sampling technology, which can reflect the distribution of noise in the first voice through the change of the amplitude of the waveform graph; wherein the waveform graph can be a time-domain graph for reflecting the time waveform of the voice signal, or a frequency spectrum graph for reflecting the amplitude-frequency distribution state of the voice signal, or a spectrogram for reflecting the characteristics of the voiceprint; and the text information corresponding to the first voice refers to the letters, punctuation marks or other symbols obtained by converting the voice signal through voice recognition technology.

[0040] In this embodiment, the display screen can be a straight screen that does not deform on the surface, such as a computer and a tablet; it can also be a flexible screen that is prone to surface deformation, such as a flexible screen mobile phone and a reader; it can also be a folding screen, such as a folding screen mobile phone, etc. One screen display interface of the folding screen mobile phone is used as the first area to display the waveform graph converted from the first voice, and the other screen display interface of the folding screen is used as the second area to display the text information transcribed from the first voice.

[0041] ​It should be noted that the folding screen refers to a foldable display screen composed of two display screens, which can change the size of the display screen by folding or unfolding, and is convenient to carry; and the flexible screen refers to a special organic light emitting display (Organic Light Emitting Display, OLED), which is a light and thin display screen that can be bent or curled, composed of polarizing plates, flexible packaging films, organic light emitting diodes, flexible films and other materials.

[0042] When receiving the input to the display screen, the display screen displays the first speech to be transcribed in response to the input.

[0043] The input can be at least one of the following ways:

[0044] First, the input can be a touch input, including but not limited to a click input, a sliding input, and a pressing input, etc.

[0045] In this embodiment, receiving the input of the user can be receiving the touch operation of the user on the display screen.

[0046] In order to reduce the user misoperation rate, the active area of the display screen can be limited to a specific area, such as a blank area of the display screen interface.

[0047] In this embodiment, after the user clicks the selection control on the display screen, a speech list containing multiple speeches is popped up, and the first speech in the speech list is selected as the first speech.

[0048] Second, the input can be a physical key input.

[0049] In this embodiment, the terminal is provided with a physical key for calling a speech list on the body of the terminal, and receiving the input of the user can be receiving the operation of pressing the corresponding physical key by the user; the input can also be a combination operation of pressing multiple physical keys at the same time.

[0050] In this embodiment, after the user clicks the physical key "#" on the display screen twice in succession, a speech list containing multiple speeches is popped up on the display screen, and the first speech in the speech list is selected as the first speech by clicking the number "1".

[0051] Third, the input can be a speech input.

[0052] In this embodiment, after the display screen is opened, a speech list containing multiple speeches is popped up on the display screen after receiving the speech "open the speech list", and then the first speech in the speech list is selected as the first speech after receiving the speech "select the first speech".

[0053] Of course, in other embodiments, the input can also be in other forms, including but not limited to character input, etc., which can be determined according to actual needs, and the present embodiment does not limit this.

[0054] When receiving the input of the first voice, the display screen displays the waveform graph corresponding to the first voice in the first area of the display screen and displays the text information corresponding to the first voice in the second area of the display screen in response to the input.

[0055] The input can be in the form of a touch input, a physical key input, a combination of the two, a voice input, and other types of input. For details, refer to the input forms described above, and the present embodiment will not be repeated here.

[0056] In this embodiment, when the input of the first voice is a touch input, for example, the conversion control corresponding to the first voice on the display screen, the display screen displays the waveform graph corresponding to the first voice in the first area and displays the text information corresponding to the first voice in the second area.

[0057] In some embodiments, when the input of the first voice is a physical key input and a combination of the two, for example, after selecting the first voice and the user clicks the physical key "*" on the display screen twice in succession, the conversion function of the first voice is triggered, and the waveform graph corresponding to the first voice is displayed in the first area and the text information corresponding to the first voice is displayed in the second area.

[0058] In some embodiments, when the input of the first voice is a voice input, for example, after receiving "convert the first voice into a waveform graph and text information", the display screen displays the waveform graph corresponding to the first voice in the first area and displays the text information corresponding to the first voice in the second area.

[0059] Step 120, receiving the first input of the user.

[0060] In this step, after displaying the waveform graph corresponding to the first voice in the first area of the display screen and displaying the text information corresponding to the first voice in the second area, the first input of the user is received, which is used to determine the decibel range, and the decibel range is used to filter out the voice of the target decibel from the first voice.

[0061] The first input can be at least one of the following:

[0062] First, the first input can be a touch input, including but not limited to click input, slide input, and press input, etc.

[0063] In this embodiment, receiving the first input of the user can be receiving the touch operation of the user on the display screen.

[0064] To reduce the user misoperation rate, the active area of the display screen can be limited to a specific area, such as the blank area of the display screen.

[0065] In this embodiment, the first input of the user is received, and the de-noising control on the display screen displays a target list containing a plurality of de-noising decibel options in response to the first input.

[0066] In this embodiment, after clicking the de-noising control on the display screen, a target list containing three de-noising options of 0dB-10dB, 0dB-20dB and 0dB-30dB is displayed, and the three de-noising options can be used to filter the sound in the corresponding decibel range.

[0067] Of course, in some embodiments, the first input of the user is received, and the numerical display area on the display screen displays a decibel value of any numerical value in response to the first input.

[0068] In this embodiment, after inputting the value 5 on the display screen, the value 5 is displayed on the numerical display area, which can be used to filter the sound of 5dB in the first voice.

[0069] In some embodiments, the first input of the user is received, and the numerical display area on the display screen displays an interval composed of any two decibel values in response to the first input.

[0070] In this embodiment, after inputting the values 5 and 35 on the display screen, the interval 5-35 is displayed on the numerical display area, which can be used to filter the sound of 5dB-30dB.

[0071] In some embodiments, when the display screen is a folding screen, the first input of the user to the display screen can be an operation on the first screen and the second screen of the folding, for example, adjusting the decibel value of the filtered noise by changing the angle between the first screen and the second screen, such as adjusting the folding angle of the first screen and the second screen to 10°, which can represent that the decibel value of the sound corresponding to the folding screen is 10dB, and also can represent that the interval of the decibel value of the sound corresponding to the folding screen is 0dB-10dB.

[0072] In some embodiments, when the display screen is a flexible screen, the first input received from the user to the display screen can be a touch operation on different areas of the flexible screen, for example, sliding the top blank area of the flexible screen can represent that the corresponding sound decibel value of the flexible screen is 10 dB, or can represent that the corresponding decibel range of the flexible screen is 0 dB-10 dB, sliding the blank area at the middle position of the flexible screen can represent that the corresponding sound decibel value of the flexible screen is 20 dB, or can represent that the corresponding decibel range of the flexible screen is 0 dB-20 dB, and so on. The first input received from the user to the display screen can also be an operation of bending a fixed angle on different areas of the flexible screen, for example, bending the top of the flexible screen by 10° represents that the corresponding sound decibel value of the flexible screen is 10 dB or the corresponding decibel range is 0 dB-10 dB, bending the top blank area at the middle position of the flexible screen by 10° represents that the corresponding sound decibel value of the flexible screen is 20 dB or the corresponding decibel range is 0 dB-20 dB, and so on.

[0073] Secondly, the first input can be a physical key input.

[0074] In this embodiment, the physical key for opening the target list is arranged on the body keyboard of the display screen, and the first input received from the user can be an operation of pressing the corresponding physical key. The first input can also be a combination operation of pressing multiple physical keys at the same time.

[0075] In this embodiment, after clicking the "OK" key on the keyboard, the target list is popped up on the display screen, and the target list containing three denoising options of 0 dB-10 dB, 0 dB-20 dB and 0 dB-30 dB is displayed. After clicking any one of the denoising options, the corresponding decibel interval can be obtained.

[0076] In some embodiments, the numerical keys on the keyboard can also be directly clicked to select the decibel value of any interval, for example, the numerical keys "5" and "35" are clicked in sequence, and "5-35" appears on the display screen, indicating that the corresponding sound decibel value of the display screen is between 5 dB and 35 dB.

[0077] Thirdly, the first input can be a voice input.

[0078] In this embodiment, after receiving the voice input of "the decibel value is 15 dB-45 dB" from the user to the display screen, the display screen displays the decibel interval of "15-45".

[0079] Of course, in other embodiments, the first input can also be in other forms, including but not limited to character input, etc., which can be determined according to actual needs, and this embodiment does not limit this.

[0080] Step 130, in response to the first input, determining a decibel range; filtering the first voice according to the decibel range, and updating the display waveform diagram and the text information.

[0081] It should be noted that the decibel range refers to a decibel value range corresponding to the noise filtered out of the first voice. The noise in the first voice refers to the noise or noise point that interferes with the normal playing of the voice. Different noise or noise points correspond to different decibel values. The required decibel range can be determined according to the plurality of decibel values obtained by the first input, and then the noise in the first voice is eliminated.

[0082] In some embodiments, the user can filter out the voice in the first voice that contains the decibel range according to the decibel range.

[0083] In this embodiment, when the user selects 0dB-20dB in the target list, it indicates that the decibel range of the filtered first voice is 0dB-20dB, and the first voice after filtering does not contain voice within 20dB. Then, the waveform diagram corresponding to the filtered first voice is updated and displayed in the first area, and the corresponding text information is updated and displayed in the second area.

[0084] In some embodiments, the user can filter out the voice in the first voice that does not include the decibel range according to the decibel range, that is, the decibel range of the filtered first voice is the decibel range.

[0085] In this embodiment, when the user selects 0dB-20dB in the target list, it indicates that the decibel range of the filtered first voice is 0dB-20dB, and the first voice after filtering only contains voice within 20dB. Then, the waveform diagram corresponding to the filtered first voice is updated and displayed in the first area, and the corresponding text information is updated and displayed in the second area.

[0086] It should be noted that the decibel range can be a fixed interval value, for example, the decibel range 0dB-10dB and 0dB-20dB, etc. The decibel range can also be an arbitrary interval value, for example, the decibel range is 8dB-62dB, etc.

[0087] In this step, after filtering the first voice according to the above decibel range, the waveform diagram and the text information corresponding to the first voice will also change. The user can evaluate the effect of filtering the first voice according to the current decibel range according to whether the waveform diagram or the text information contains invalid waveform or invalid character.

[0088] According to the method provided in the embodiment of the present application, the waveform graph and the transcribed text information of the first voice are displayed on two interfaces respectively, and the noise of the waveform graph and the text information is filtered by adjusting the sound decibel range to update the first voice, so that the effective information in the first voice can be quickly obtained, and the waveform graph and the text information are updated according to the determined decibel range on the two interfaces at the same time, the filtering result can be intuitively and synchronously displayed to the user, the use is convenient, and the operation difficulty of the user is greatly reduced.

[0089] In some embodiments, the step 110 of displaying, based on the first voice, a waveform graph corresponding to the first voice on a first area of the display screen and displaying text information obtained by transcribing the first voice on a second area of the display screen comprises: performing semantic segmentation on the first voice to obtain a plurality of voice segments; and based on the plurality of voice segments, displaying corresponding multi-segment waveform graphs on the first area and displaying corresponding multi-segment text information on the second area.

[0090] In this embodiment, the first voice is subjected to semantic segmentation to obtain a plurality of voice segments, each voice segment is then converted into a waveform graph by a voice sampling unit built in the display screen and displayed on the first area of the display screen, and each voice segment is converted into corresponding text information by a voice recognition unit and displayed on the second area of the display screen.

[0091] In this embodiment, the display screen is built with a voice sampling unit and a voice recognition unit, the voice sampling unit can convert one or more voices into a corresponding number of waveform graphs, and the voice recognition unit can transcribe one or more voices into a corresponding number of text information.

[0092] In Figure 2 In the embodiment shown in the figure, after the first voice stored in the display screen 210 is segmented into a plurality of voice segments, each voice segment is converted into a corresponding number of waveform graphs 2111 by a voice sampling unit and displayed on the first area 211, and each voice segment is converted into a corresponding number of text information 2121 by a voice recognition unit and displayed on the second area 212.

[0093] Of course, in some embodiments, the first voice can not be segmented into a plurality of voice segments, but can be directly converted into a waveform graph and text information and displayed on the corresponding area.

[0094] According to the method provided in the embodiment of the present application, after the first voice is segmented into a plurality of voice segments, the voice segments are displayed in two areas of the display screen in the form of waveform graphs and text information respectively, which is helpful for visual analysis of the noise contained in each voice segment and provides convenience for subsequent dynamic updating of the waveform graph and the text information.

[0095] Next, how to update the waveform chart and text information is described.

[0096] In some embodiments, after step 110, displaying the corresponding multi-segment waveform chart in the first area and displaying the corresponding multi-segment text information in the second area based on the plurality of voice segments, the method further comprises: receiving a second input of a user on a target waveform chart in the multi-segment waveform chart; and playing the voice corresponding to the target waveform chart in response to the second input.

[0097] In this embodiment, one voice is divided into a plurality of voice segments, the plurality of voice segments are converted into a corresponding number of waveform charts by a voice sampling unit, and the plurality of waveform charts are displayed in the first area; the waveform charts corresponding to the plurality of voice segments are displayed in the second area by a voice recognition unit.

[0098] It should be noted that the target waveform chart is one or more of the multi-segment waveform chart, and the target waveform chart is the waveform chart to be updated.

[0099] In this embodiment, the second input of the user is received, and a playing interface of the target waveform chart is displayed in response to the target waveform chart in the plurality of waveform charts, and the voice corresponding to the waveform chart is played by a microphone built in the display screen.

[0100] The second input can be at least one of the following ways:

[0101] Firstly, the second input can be a touch input, including but not limited to a click input, a sliding input, and a pressing input, etc.

[0102] In this implementation, the second input of the user can be a touch operation of the user on the target waveform chart.

[0103] In order to reduce the user misoperation rate, the action area of the target waveform chart in the second area can be limited in a specific area, such as a blank area of the target waveform chart.

[0104] In this embodiment, long pressing the area where the first segment waveform chart displayed in the first area of the display screen is located can play the voice corresponding to the waveform chart.

[0105] Secondly, the second input can be a physical key input.

[0106] In this implementation, the body of the display screen is provided with a physical key for playing the voice corresponding to the target waveform chart, and the second input of the user can be an operation of pressing the corresponding physical key by the user; the second input can also be a combination operation of pressing a plurality of physical keys at the same time.

[0107] In this embodiment, the user can play the voice corresponding to the first waveform graph by clicking the entity keys corresponding to "*" and "1" respectively, play the voice corresponding to the first waveform graph by clicking the entity keys corresponding to "*" and "2" respectively, and so on. This embodiment will not be described again.

[0108] Thirdly, the second input can be a voice input.

[0109] In this implementation, after displaying the multiple waveform graphs in the second area of the display screen, the voice corresponding to the first waveform graph can be played after receiving the voice "play voice segment 1".

[0110] Of course, in other embodiments, the second input can also be in other forms, including but not limited to character input, etc., which can be determined according to actual needs, and this embodiment does not limit this.

[0111] According to the method provided in the embodiments of the present application, the second input to the target waveform graph is received to play the voice of the target waveform graph, which helps to accurately identify the noise in the voice corresponding to each segmented waveform graph, and the operation is simple.

[0112] Next, taking the voice transcription by the above method as an example, the description is as follows.

[0113] I. Voice transcription based on folding screen

[0114] In some embodiments, the display screen is a folding screen, and includes a foldable first screen and a second screen, the first area is located in the first screen, and the second area is located in the second screen; the first input includes: folding the first screen and the second screen; in response to the first input, determining the decibel range includes: determining the decibel range based on the angle between the first screen and the second screen.

[0115] In this embodiment, the display screen is a folding screen, and the first screen and the second screen of the folding screen are located in the first area and the second area respectively. The first input of the user is received, and the display screen determines the decibel range in response to the first input. It can be determined by changing the angle between the first screen and the second screen to determine the decibel range of the filtered first voice, and filtering the first voice according to the decibel range, and finally updating the waveform graph in the first area and the text information in the second area.

[0116] In this embodiment, the decibel range is changed by changing the included angle of the first screen and the second screen of the folding screen, for example, the required decibel range is determined according to the mapping relationship established between the change of the included angle of the two screens in the folding screen and different decibel ranges. The mapping relationship can be that when the change range of the included angle of the two screens in the folding screen is 0°-10°, the corresponding decibel range of the folding screen is set to 0°-10°, at this time, the folding screen can filter the 0dB-10dB sound in the first voice, or only keep the 0dB-10dB sound in the first voice; when the change range of the included angle of the two screens in the folding screen is 10°-20°, the corresponding decibel range of the folding screen is set to 0°-20°, at this time, the folding screen can filter the 0dB-20dB sound in the first voice, or only keep the 0dB-20dB sound in the first voice, and so on.

[0117] It should be noted that the embodiment can control the decibel range according to the change of the included angle of the two screens in the folding screen, and when the included angle of the two screens increases by 1°, the corresponding decibel range increases by 1dB, and when the included angle of the two screens decreases by 1°, the corresponding decibel range decreases by 1dB.

[0118] In Figures 2-4 In the embodiment shown in FIG. 21, after the first voice is segmented into a plurality of voice segments, the waveform diagram 2111 and the text information 2121 corresponding to the plurality of voice segments are displayed on the first area 211 and the second area 212 of the display screen 210 respectively, and the decibel range of the first voice is determined by changing the included angle of the first screen and the second screen of the folding screen. Finally, the waveform diagram and the text information corresponding to the filtered first voice are updated and displayed on the first area 211 and the second area 212 of the display screen 210 respectively.

[0119] In Figure 3 In the embodiment shown in FIG. 21, when the included angle between the first area 211 and the second area 212 of the display screen 210 increases by 10°, the display screen 210 can filter the voice within 10dB in the first voice, and update and display the waveform diagram 2112 corresponding to the filtered first voice on the first area 211 of the display screen 210, and update and display the text information 2122 corresponding to the filtered first voice on the second area 212 of the display screen 210.

[0120] In some embodiments, when the voice corresponding to the updated waveform diagram is played or the updated text information is observed, the above-mentioned invalid waveform or invalid character still exists, and the decibel range of the folding screen needs to be further adjusted to filter the first voice.

[0121] In Figure 4In the illustrated embodiment, when the angle between the first region 211 and the second region 212 of the display screen 210 is increased by 20°, the display screen 210 can filter the voice in the first voice within 20 dB, and update the display of the waveform graph 2113 corresponding to the filtered first voice in the first region 211 of the display screen 210, and update the display of the text information 2123 corresponding to the filtered first voice in the second region 212.

[0122] When the voice corresponding to the updated waveform graph is played or the updated text information is observed, and there is no invalid waveform or invalid character as described above, it can be considered that the filtering process of the first voice has ended.

[0123] According to the method provided in the embodiments of the present application, when the display screen is a folding screen, the first input of the user is received as an operation of changing the angle between the two screens of the folding screen, the decibel range for denoising the first voice is determined, the operation of denoising the voice with the folding screen is simplified, and after denoising the first voice according to the decibel range, the waveform graph and the text information corresponding to the voice are respectively updated and displayed in the first screen and the second screen of the folding screen, which is more intuitive.

[0124] II. Voice transcription based on a flexible screen.

[0125] In some embodiments, the display screen is a flexible screen; the first input includes an operation of folding a target region of the flexible screen; and in response to the first input, determining the decibel range includes determining the decibel range based on the position of the target region or determining the decibel range based on the angle at which the target region is folded.

[0126] In this embodiment, the display screen is a flexible screen, the first input of the user is received, the target region determines the decibel range in response to the first input, the decibel range for filtering the first voice can be determined by receiving different positions of the target region of the first input, and the first voice is filtered according to the decibel range, and finally the waveform graph in the first region and the text information in the second region are updated to update the waveform graph in the first region and the text information in the second region, wherein the first region is used to display the waveform graph corresponding to the first voice, and the second region is used to display the text information corresponding to the first voice.

[0127] In this embodiment, the required decibel range can be determined according to the mapping relationship between the operation of the target region of the flexible screen and different decibel ranges; such a mapping relationship can be that when there are multiple target regions, different decibel ranges can be determined according to different positions of the target region receiving the first input. For example, when the first target region receives the first input, the corresponding decibel range of the flexible screen is 0 dB-10 dB, when the second target region receives the first input, the corresponding decibel range of the flexible screen is 0 dB-20 dB, and so on.

[0128] InFigure 5 In the embodiment shown, the display screen 210 is a flexible screen. The display screen 210 has three target areas 213 (corresponding to area ①, area ② and area ③). After clicking on the target area 213 (corresponding to ①), the display screen 210 can filter out the first speech within 10dB, and update the waveform diagram 2112 corresponding to the filtered first speech in the first area 211 of the display screen 210, and update the text information 2122 corresponding to the filtered first speech in the second area 212.

[0129] In some embodiments, when there is only one target area, different decibel ranges can be determined based on the duration or frequency of the target area. For example, if the target area is pressed for more than 2 seconds, the decibel range corresponding to the folding screen is 0dB-10dB; if the target area is pressed for more than 4 seconds, the decibel range corresponding to the folding screen is 0dB-20dB. The first speech is filtered based on the decibel range, and then the waveform and text information corresponding to the filtered first speech are updated and displayed in the first and second areas, respectively.

[0130] In some embodiments, the decibel range can also be determined by folding the target area of ​​the flexible screen.

[0131] In this embodiment, when there are multiple target areas, different decibel ranges can be determined by folding multiple areas at a fixed angle, which can be between 0° and 90°.

[0132] In this embodiment, there are three target areas on the flexible screen. When the first target area is folded by 10°, the corresponding decibel range of the flexible screen is 0dB-10dB. When the second target area is folded by 10°, the corresponding decibel range of the flexible screen is 0dB-20dB. When the third target area is folded by 10°, the corresponding decibel range of the flexible screen is 0dB-30dB.

[0133] In some embodiments, when there is only one target area, different decibel ranges can be determined by folding the target area at different angles.

[0134] In this embodiment, when the target area on the flexible screen is folded by 10°, the corresponding decibel range of the flexible screen is 0dB-10dB; when the target area is folded by 20°, the corresponding decibel range of the flexible screen is 0dB-20dB; and when the target area is folded by 30°, the corresponding decibel range of the flexible screen is 0dB-30dB.

[0135] In some embodiments, if the above-mentioned invalid waveforms or invalid characters still exist after playing the voice corresponding to the updated waveform or observing the updated text information, it is necessary to continue to adjust the decibel range of the flexible screen to filter the first voice, such as the method of adjusting the decibel range of the foldable screen to filter the first voice as described above. This embodiment will not repeat the details.

[0136] According to the method provided in the embodiments of the present application, when the display screen is a flexible screen, the decibel range of the first voice denoising is determined by receiving the operation of the user on the target area, which brings convenience to the voice denoising process of the flexible screen, and the waveform graph and the text information corresponding to the voice are respectively updated in the first area and the second area of the flexible screen after the first voice denoising according to the decibel range, which is simple to operate and more intuitive to display.

[0137] The execution subject of the voice transcription method provided in the embodiments of the present application can be a voice transcription device. The voice transcription device provided in the embodiments of the present application is described by taking the voice transcription device as an example.

[0138] The embodiments of the present application further provide a voice transcription device.

[0139] As shown in Figure 6 The voice transcription device includes a display module 610, a first receiving module 620 and a first processing module 630.

[0140] The display module 610 is configured to display a waveform graph corresponding to the first voice in a first area of a display screen and display text information obtained by transcribing the first voice in a second area of the display screen based on the first voice.

[0141] The first receiving module 620 is configured to receive a first input of a user.

[0142] The first processing module 630 is configured to determine a decibel range in response to the first input.

[0143] The first voice is filtered according to the decibel range, and the waveform graph and the text information are updated.

[0144] According to the voice transcription device provided in the embodiments of the present application, the waveform graph of the first voice and the transcribed text information are respectively displayed in two interfaces by the display module 610, and after receiving the first input of the user by the first receiving module 620, the noise of the waveform graph and the text information is filtered by adjusting the decibel range by the first processing module 630 in response to the first input to update the first voice, which is more intuitive to display and greatly reduces the operation difficulty of the user.

[0145] In some embodiments, the device further includes:

[0146] The second processing module is configured to perform semantic segmentation on the first voice to obtain a plurality of voice segments; and the display module is further configured to display a plurality of corresponding waveform graphs in the first area and display a plurality of corresponding transcribed text information in the second area based on the plurality of voice segments.

[0147] According to the voice transcription device provided in the embodiment of the present application, after the first voice is segmented into a plurality of voice segments, the voice segments are displayed in two areas of the display screen in the display modes of waveform graphs and text information, which helps to visually analyze the noise contained in each voice segment and provides convenience for dynamically updating the waveform graphs and the text information in the subsequent process.

[0148] In some embodiments, after the corresponding plurality of segmented waveform graphs are displayed in the first area and the corresponding plurality of segmented text information is displayed in the second area based on the plurality of voice segments, the device further comprises:

[0149] The second receiving module is configured to receive a second input of a target waveform graph in the plurality of segmented waveform graphs by a user.

[0150] The third processing module is configured to play the voice corresponding to the target waveform graph in response to the second input.

[0151] According to the voice transcription device provided in the embodiment of the present application, the voice of the target waveform graph is played by receiving the second input of the target waveform graph, which helps to accurately identify the noise in the voice corresponding to each segmented waveform graph and is simple to operate.

[0152] In some embodiments, the display screen is a foldable screen and includes a first screen and a second screen that can be folded, the first area is located in the first screen, and the second area is located in the second screen.

[0153] The first input includes an operation of folding the first screen and the second screen.

[0154] In response to the first input, the decibel range is determined, including:

[0155] The first processing module is further configured to determine the decibel range based on the included angle of the first screen and the second screen.

[0156] According to the voice transcription device provided in the embodiment of the present application, when the display screen is a foldable screen, the decibel range for de-noising the first voice is determined by receiving an operation of a user changing the included angle between the two screens of the foldable screen, which simplifies the operation of de-noising the voice by using the foldable screen, and the waveform graph and the text information corresponding to the voice are respectively updated and displayed in the first screen and the second screen of the foldable screen after de-noising the voice according to the decibel range, which is more intuitive. In some embodiments, the display screen is a flexible screen.

[0157] The first input includes an operation of folding a target area of the flexible screen.

[0158] In response to the first input, the decibel range is determined, including:

[0159] The first processing module is further configured to determine the decibel range based on the position of the target area.

[0160] Alternatively, the first processing module is further configured to determine the decibel range based on an angle at which the target region is folded.

[0161] According to the voice transcription apparatus provided in the embodiments of the present application, when the display screen is a flexible screen, the decibel range for the first voice denoising is determined by receiving the operation of the target region by the user, which brings convenience to the voice denoising process of the flexible screen, and the waveform graph and the text information corresponding to the voice are respectively updated in the first region and the second region of the flexible screen after the first voice denoising according to the decibel range, which is simple to operate and more intuitive to display.

[0162] The voice transcription apparatus in the embodiments of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than the terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and the like, and can also be a Network Attached Storage (NAS), a personal computer (PC), or a self-service machine, and the like, and the embodiments of the present application are not limited specifically.

[0163] The voice transcription apparatus in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an IOS operating system, or other possible operating systems, and the embodiments of the present application are not limited specifically.

[0164] The voice transcription apparatus provided in the embodiments of the present application can implement Figures 1 to 5 The method embodiments implement various processes, and to avoid repetition, the various processes are not described herein again.

[0165] Optionally, as shown in Figure 7 The embodiments of the present application further provide an electronic device 700, which includes a processor 701, a memory 702, a program or instructions stored in the memory 702 and executable on the processor 701. The program or instructions are executed by the processor 701 to implement various processes of the voice transcription method embodiments described above, and achieve the same technical effects. To avoid repetition, the various processes are not described herein again.

[0166] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0167] Figure 8 A hardware structure schematic diagram of an electronic device according to an embodiment of the present application.

[0168] The electronic device 800 includes, but is not limited to, a radio frequency unit 801, a network module 802, an audio output unit 803, an input unit 804, a sensor 805, a display unit 806, a user input unit 807, an interface unit 808, a memory 809, and a processor 810, etc.

[0169] Those skilled in the art can understand that the electronic device 800 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 810 through a power management system, so as to realize the functions of power management, such as charging, discharging, and power consumption management, through the power management system. Figure 8 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.

[0170] The display unit 806 is configured to display a waveform graph corresponding to the first voice in a first area of the display screen and display text information obtained by transcribing the first voice in a second area of the display screen based on the first voice.

[0171] The user input unit 807 is configured to receive a first input of a user.

[0172] The processor 810 is configured to determine a decibel range in response to the first input.

[0173] The first voice is filtered according to the decibel range, and the waveform graph and the text information are updated.

[0174] According to the electronic device provided by the embodiments of the present application, the waveform graph of the first voice and the transcribed text information are respectively displayed in two interfaces, which can intuitively and synchronously show the filtering result to the user, is convenient to use, and adjusts the decibel range to filter the noise of the waveform graph and the text information to update the first voice, which greatly reduces the operation difficulty of the user. Optionally, the processor 810 is further configured to perform semantic segmentation on the first voice to obtain a plurality of voice segments.

[0175] The display unit 806 is further configured to display a plurality of waveform graphs corresponding to the plurality of voice segments in the first area and display a plurality of text information corresponding to the plurality of voice segments in the second area based on the plurality of voice segments.

[0176] According to the electronic device provided in the embodiment of the present application, after the first voice is segmented into a plurality of voice segments, the voice segments are displayed in two areas of the display screen in the display modes of waveform graphs and text information, which is helpful for visual analysis of the noise contained in each voice segment and provides convenience for dynamic updating of the waveform graphs and the text information in the subsequent process.

[0177] Optionally, the user input unit 807 is configured to receive a second input of a target waveform graph in the plurality of segmented waveform graphs.

[0178] The audio output unit 8038 is configured to play the voice corresponding to the target waveform graph in response to the second input.

[0179] According to the electronic device provided in the embodiment of the present application, the voice of the target waveform graph is played by receiving the second input of the target waveform graph, which is helpful for accurate identification of the noise in the voice corresponding to each segmented waveform graph and simple operation.

[0180] Optionally, the display screen is a foldable screen and includes a first screen and a second screen that can be folded, the first area is located on the first screen, and the second area is located on the second screen; the first input includes an operation of folding the first screen and the second screen; and in response to the first input, the decibel range is determined, including:

[0181] The processor 810 is further configured to determine the decibel range based on an included angle between the first screen and the second screen.

[0182] According to the electronic device provided in the embodiment of the present application, when the display screen is a foldable screen, the operation of voice denoising by using the foldable screen is simplified by receiving an operation of the user changing an included angle between two screens in the foldable screen to determine the decibel range for denoising the first voice, and the waveform graph and the text information corresponding to the voice are respectively updated and displayed on the first screen and the second screen of the foldable screen after the first voice is denoised according to the decibel range, and the display is more intuitive.

[0183] Optionally, the display screen is a flexible screen; and the first input includes an operation of folding a target area of the flexible screen.

[0184] In response to the first input, the decibel range is determined, including:

[0185] The processor 810 is further configured to determine the decibel range based on a position of the target area.

[0186] Alternatively, the processor 810 is further configured to determine the decibel range based on an angle at which the target area is folded.

[0187] According to the electronic device provided in the embodiments of the present application, when the display screen is a flexible screen, the decibel range of the first voice de-noising is determined by receiving the operation of the target area by the user, which brings convenience to the voice de-noising process of the flexible screen, and the waveform graph and the text information corresponding to the voice are respectively updated in the first area and the second area of the flexible screen after the first voice de-noising according to the decibel range, which is simple to operate and more intuitive to display.

[0188] It should be understood that in the embodiments of the present application, the input unit 804 can include a graphics processor (GPU) 8041 and a microphone 8042. The graphics processor 8041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 806 can include a display panel 8061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 can include two parts of a touch detection device and a touch controller. The other input devices 8072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, an operating rod, etc., which will not be described here.

[0189] The memory 809 can be used to store software programs and various data. The memory 809 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 809 can include a volatile memory or a non-volatile memory, or the memory 809 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 809 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0190] The processor 810 can include one or more processing units; optionally, the processor 810 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 810.

[0191] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned voice conversion method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0192] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0193] The chip provided by the embodiment of the present application includes a processor and a communication interface, the communication interface is coupled with the processor, the processor is used to run programs or instructions, realizes the processes of the voice conversion method embodiments, and achieves the same technical effects. To avoid repetition, details are not described here.

[0194] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip chip, etc.

[0195] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in reverse order, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0196] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk, etc.), including a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network equipment, etc.) execute the method described in each embodiment of the present application.

[0197] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.

Claims

1. A voice transcription method, characterized by, The method comprises: displaying, based on a first voice, a waveform graph corresponding to the first voice in a first area of a display screen and displaying text information obtained by transcribing the first voice in a second area of the display screen; receiving a first input of a user, the first input being obtained by the user performing visual analysis on noise in the voice based on the waveform graph and the text information; determining, in response to the first input, a decibel range, the decibel range being used to screen or filter out voice of a target decibel from the first voice; filtering the first voice according to the decibel range and updating the display of the waveform graph and the text information. 2.The voice transcription method of claim 1, wherein, The method comprises: performing semantic segmentation on the first voice to obtain a plurality of voice segments; based on the plurality of voice segments, displaying corresponding multi-segment waveform graphs in the first area and displaying corresponding multi-segment text information obtained by transcribing in the second area. 3.The voice transcription method of claim 2, wherein, After the above step, the method further comprises: receiving a second input of the user on a target waveform graph in the multi-segment waveform graphs; playing the voice corresponding to the target waveform graph in response to the second input.

4. The voice transcription method of any one of claims 1-3, wherein, The display screen is a foldable screen and comprises a first foldable screen and a second foldable screen, the first area is located on the first screen, and the second area is located on the second screen. The first input comprises an operation of folding the first screen and the second screen. The method comprises: determining the decibel range based on the included angle between the first screen and the second screen.

5. The voice transcription method according to any one of claims 1-3, wherein, The display screen is a flexible screen. The first input comprises an operation of folding a target area of the flexible screen. The method comprises: determining the decibel range based on the position of the target area. Alternatively, determining the decibel range based on the angle at which the target area is folded.

6. A speech transcription apparatus characterized by comprising: The method comprises: a display module configured to display, based on a first voice, a waveform graph corresponding to the first voice in a first area of a display screen and display text information obtained by transcribing the first voice in a second area of the display screen; a first receiving module configured to receive a first input of a user, the first input being obtained by the user performing visual analysis on noise in the voice based on the waveform graph and the text information; a first processing module configured to determine, in response to the first input, a decibel range, the decibel range being used to screen or filter out voice of a target decibel from the first voice; filtering the first voice according to the decibel range and updating the display of the waveform graph and the text information.

7. The speech transcription apparatus according to claim 6, wherein The device further comprises: a second processing module configured to perform semantic segmentation on the first voice to obtain a plurality of voice segments; the display module is further configured to display, based on the plurality of voice segments, corresponding multi-segment waveform graphs in the first area and display corresponding multi-segment text information obtained by transcribing in the second area.

8. The speech transcription apparatus according to claim 7, wherein After the display of the corresponding multi-segment waveform graph in the first area and the corresponding multi-segment text information in the second area based on the plurality of voice segments, the device further comprises: a second receiving module configured to receive a second input of a target waveform graph in the multi-segment waveform graph; a third processing module configured to play the voice corresponding to the target waveform graph in response to the second input.

9. The speech transcription apparatus according to any one of claims 6 to 8, characterized by, The display screen is a foldable screen, and comprises a first foldable screen and a second foldable screen, the first area is located on the first screen, and the second area is located on the second screen. The first input comprises an operation of folding the first screen and the second screen. The determination of the decibel range in response to the first input comprises: The first processing module is further configured to determine the decibel range based on the angle between the first screen and the second screen.

10. The speech transcription apparatus according to any one of claims 6 to 8, wherein The display screen is a flexible screen. The first input comprises an operation of folding a target area of the flexible screen. The determination of the decibel range in response to the first input comprises: The first processing module is further configured to determine the decibel range based on the position of the target area. Alternatively, the first processing module is further configured to determine the decibel range based on the angle at which the target area is folded.

Citation Information

Patent Citations

  • A screen shooting method and terminal device

    CN109542306A

  • Audio processing method and device, mobile terminal and computer readable storage medium

    CN110349594A

  • Audio processing method and electronic equipment

    CN111445927A

  • Voice processing method and electronic equipment

    CN113205815A