Interview conversation audio processing method and device, electronic equipment and storage medium
By using a pre-trained language model to extract job-related vocabulary during the interview process to generate an intervention hot word list, and combining it with a speech recognition system and error correction instructions, the problem of inaccurate recognition of professional terms in the interview-converted text was solved, thus improving the accuracy and efficiency of the converted text.
Patent Information
- Application Number
- CN202411333025.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-09-24
AI Technical Summary
Existing automatic speech recognition technology is not accurate enough in recognizing professional or industry-specific terms during interviews, which leads to the need for manual proofreading and modification of the converted text, increasing time and manpower costs.
By using a pre-trained language model, common and non-common words related to the job are extracted from the resumes of interview candidates and the standard answers to interview questions, an intervention hot word list is generated, and the interview dialogue audio is converted into text based on the list. The accuracy of the converted text is improved by combining a speech recognition system and preset error correction instructions.
By generating a job-related hot word list, errors in recognizing professional terms and industry jargon were reduced, the accuracy of converting interview dialogue audio to text was improved, and the need for manual proofreading was reduced.
Smart Images

Figure CN119274558B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI interview technology, and more specifically, to an audio processing method, apparatus, electronic device, and storage medium for interview dialogues. Background Technology
[0002] In traditional interview processes, recording interview conversations typically relies on manual note-taking or audio recording. This method is not only time-consuming but also prone to inaccuracies. In recent years, automatic speech recognition technology has been introduced into the transcription of interview conversations, automatically converting audio into text. However, existing automatic speech recognition technology still falls short of ideal accuracy when dealing with technical terms or field-specific interview questions, especially when candidates use specialized or industry-specific terms, leading to frequent recognition errors. Consequently, the converted text often requires manual proofreading and revision, further increasing time and labor costs. Summary of the Invention
[0003] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing interview dialogue audio, which can improve the accuracy of the converted text corresponding to the interview dialogue audio.
[0004] The technical solution of this invention can be implemented as follows:
[0005] In a first aspect, the present invention provides a method for processing interview dialogue audio, the method comprising:
[0006] Using a pre-trained language model, the first and second vocabulary words are extracted from the resumes of interview candidates and the standard answers to interview questions, respectively. The first and second vocabulary words include common words and non-everyday words related to the job position applied for by the interview candidates.
[0007] Based on the first and second vocabulary words, generate an intervention hot word list;
[0008] Based on the aforementioned hot word list, the interview dialogue audio is converted into text to obtain the converted text corresponding to the interview dialogue audio.
[0009] Optionally, the step of generating an intervention hot word list based on the first and second words includes:
[0010] If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed the preset upper limit, then each word in the first vocabulary and the second vocabulary will be added as a hot word to the intervention hot word list.
[0011] If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds the preset upper limit, then each word in the first vocabulary will be added as a hot word to the intervention hot word list.
[0012] Optionally, the electronic device is equipped with a speech recognition system, and the step of performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio includes:
[0013] The interview dialogue audio and the intervention hot word list are input into the speech recognition system, so that the speech recognition system performs speech recognition on the interview dialogue audio according to each hot word in the intervention hot word list and outputs the recognized text corresponding to the interview dialogue audio.
[0014] The identified text is used as the converted text.
[0015] Optionally, the step of the speech recognition system performing speech recognition on the interview dialogue audio based on each hot word in the intervention hot word list and outputting the recognized text corresponding to the interview dialogue audio includes:
[0016] Before performing speech recognition, the speech recognition system assigns a recognition weight to each hot word based on its frequency of occurrence in the intervention hot word list.
[0017] The speech recognition system performs speech recognition on the interview dialogue audio according to the recognition weight of each hot word, and dynamically adjusts the recognition weight of each hot word in real time according to the recognition situation during the speech recognition process, until the speech recognition is completed and the recognized text is output.
[0018] Optionally, the step of dynamically adjusting the recognition weight of each hot word in real time based on the recognition results during the speech recognition process includes:
[0019] For each hot word, if the speech recognition system detects that the hot word appears in the recognized text generated up to the current time, the recognition weight of the hot word is increased.
[0020] Optionally, the step of dynamically adjusting the recognition weight of each hot word in real time based on the recognition results during the speech recognition process includes:
[0021] For each of the hot words, if the speech recognition system detects that the phoneme of the hot word matches a segment of the interview dialogue audio being recognized at the current moment, the recognition weight of the hot word is increased.
[0022] Optionally, before using the identified text as the converted text, the step of performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio further includes:
[0023] Using the pre-trained language model, the recognized text is corrected according to preset error correction instructions, and the corrected recognized text is used as the converted text.
[0024] Secondly, the present invention provides an audio processing apparatus for interview dialogues, the apparatus comprising:
[0025] The extraction module is used to extract the first vocabulary and the second vocabulary from the interview candidate's resume and the standard answer to the interview question, respectively, using a pre-trained language model. The first vocabulary and the second vocabulary include common words and non-everyday words related to the position the interview candidate is applying for.
[0026] A generation module is used to generate an intervention hot word list based on the first word and the second word;
[0027] The conversion module is used to perform text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio.
[0028] Thirdly, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the interview dialogue audio processing method as described in the first aspect above.
[0029] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the interview dialogue audio processing method described in the first aspect above.
[0030] Compared to existing technologies, the interview dialogue audio processing method provided by this invention involves: using a pre-trained language model to extract a first vocabulary and a second vocabulary, including commonly used words and non-everyday words related to the interviewee's position, from the interviewee's resume and the standard answers to the interview questions; generating an intervention hot word list based on the first and second vocabulary; and performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio. Because this invention extracts non-everyday words and commonly used words related to the interviewee's position from the interviewee's resume and the standard answers to the interview questions to generate an intervention hot word list for text conversion processing of the interview dialogue audio, it reduces recognition errors and improves the accuracy of the obtained converted text when the interviewee uses professional terms or industry jargon.
[0031] Other features and advantages disclosed in this invention will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques disclosed in this invention.
[0032] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 A flowchart illustrating an audio processing method for an interview dialogue provided in an embodiment of the present invention;
[0035] Figure 2 A flowchart illustrating the implementation process of step S103 provided in an embodiment of the present invention. Figure 1 ;
[0036] Figure 3 A flowchart illustrating the implementation process of step S103 provided in an embodiment of the present invention. Figure 2 ;
[0037] Figure 4 A functional unit block diagram of an interview dialogue audio processing device provided in an embodiment of the present invention;
[0038] Figure 5 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention.
[0039] Icons: 100 - Interview dialogue audio processing device; 101 - Extraction module; 102 - Generation module; 103 - Conversion module; 200 - Electronic device; 210 - Memory; 220 - Processor. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0041] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0042] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0043] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0044] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.
[0045] To improve the accuracy of converting interview dialogue audio into text, this invention provides an interview dialogue audio processing method, which will be described in detail below.
[0046] Please refer to Figure 1 The interview dialogue audio processing method includes steps S101 to S103.
[0047] S101 uses a pre-trained language model to extract the first and second words from the resumes of interview candidates and the standard answers to interview questions, respectively.
[0048] Among them, the first and second vocabulary words are common words and non-everyday words related to the position the interview candidate is applying for.
[0049] The first vocabulary includes commonly used words related to the position the interviewee is applying for. These words can be campus terms, company terms, professional terms, and industry terms extracted from information such as "professional skills," "work experience," "project experience," and "internship experience" in the interviewee's resume.
[0050] The candidate's resume is input into a pre-trained language model, and appropriate prompts are written so that the pre-trained language model can extract the first word from the candidate's resume according to the prompts.
[0051] For example, when extracting the first vocabulary, the prompts written and input into the pre-trained language model could be: "Please identify and list all education-related terms and phrases from the following resume text, such as degrees, school names, and educational programs," "Please extract all company-specific terms and phrases from this resume, including job titles, department names, and corporate structure," "Please identify and list all industry-specific terms and concepts in the resume that reflect the industry and market in which the individual's experience is located," "Please extract a list of professional and technical terms from the resume that are not commonly used in everyday language and focus on terms specific to a particular field," and "Please extract and categorize the following types of terms from the resume: 1. Campus terms (e.g., university, degree, college); 2. Company terms (e.g., CEO, department, merger); 3. Professional terms (e.g., software engineer, financial analyst); 4. Industry terms (e.g., healthcare, automotive, technology); 5. Non-everyday terms (e.g., quantum computing, bioremediation); Please provide structured output with terms clearly listed for each category."
[0052] Similarly, the second vocabulary includes commonly used words related to the position the interviewee is applying for, which may be professional or industry-specific terms extracted from the standard answers to interview questions.
[0053] The standard answers to interview questions can be pre-set by the question setter or generated by using a pre-trained language model to answer the interview questions.
[0054] The interview questions and their corresponding standard answers are input into a pre-trained language model, and appropriate prompts are written so that the pre-trained language model can extract the second vocabulary from the interview questions and standard answers according to the prompts.
[0055] For example, when extracting the second vocabulary, the prompts written and input into the pre-trained language model could be: "Please identify all industry-specific terms and concepts in the input text that reflect the industry and market in which you have personal experience," "Please extract a list of technical terms and jargon from the input text that are not commonly used in everyday language and focus on terms specific to a particular field," or "Please extract and categorize the following types of terms from the input text: 1. Technical terms (e.g., software engineer, financial analyst); 2. Industry terms (e.g., healthcare, automotive, technology); 3. Non-everyday terms (e.g., quantum computing, bioremediation); Please provide structured output with terms clearly listed for each category."
[0056] S102, Generate an intervention hot word list based on the first and second words.
[0057] The intervention hot word list has a capacity limit, which is the maximum number of hot words that the intervention hot word list can contain. The intervention hot word list used to convert audio into text is generated based on the number of words in the first vocabulary, the number of words in the second vocabulary, and the capacity limit of the intervention hot word list.
[0058] In possible implementations, step S102 can have the following two scenarios:
[0059] In scenario one, if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed the preset upper limit, then each word in the first and second vocabulary will be added as a hot word to the intervention hot word list.
[0060] Scenario 2: If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds the preset upper limit, then each word in the first vocabulary will be added as a hot word to the intervention hot word list.
[0061] The preset upper limit refers to the maximum capacity of the intervention hot word list, meaning that the number of hot words contained in the intervention hot word list cannot exceed the preset upper limit.
[0062] Understandably, if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed the capacity limit of the intervention hot word list, then all words in the first vocabulary and the second vocabulary can be added as hot words to the intervention hot word list.
[0063] If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds the capacity limit of the intervention hot word list, then the first vocabulary has a higher priority than the second vocabulary, that is, each word in the first vocabulary is added to the intervention hot word list as a priority.
[0064] The reason for this is that the first word is extracted from the candidate's resume, which is written by the candidate. Consequently, the candidate will refer to their own resume when answering questions. In other words, the first word is more closely related to the candidate and is more likely to be mentioned by the candidate during the interview.
[0065] S103, based on the intervention hot word list, perform text conversion processing on the interview dialogue audio to obtain the converted text corresponding to the interview dialogue audio.
[0066] In a possible implementation, a voice recognition system is deployed on the electronic device, and step S103 may include... Figure 2 The sub-steps S103-1 to S103-2 are shown.
[0067] S103-1, Input the interview dialogue audio and the intervention hot word list into the speech recognition system, so that the speech recognition system can perform speech recognition on the interview dialogue audio according to each hot word in the intervention hot word list, and output the recognized text corresponding to the interview dialogue audio.
[0068] A typical speech recognition system includes components such as "sound capture," "preprocessing," "feature extraction," "acoustic model," "language model," "decoder," and "post-processing." By integrating an intervention hot word list into the acoustic model, language model, and decoder, the speech recognition system's performance in recognizing professional terms or industry jargon can be improved, thereby meeting the needs of speech-to-text conversion in interview scenarios.
[0069] The process by which a speech recognition system performs speech recognition on the interview dialogue audio based on each hot word in the intervention hot word list and outputs the corresponding recognized text can be as follows:
[0070] Before performing speech recognition, the speech recognition system assigns a recognition weight to each hot word based on its frequency of occurrence in the intervention hot word list.
[0071] The speech recognition system performs speech recognition on the interview dialogue audio based on the recognition weight of each hot word, and dynamically adjusts the recognition weight of each hot word in real time according to the recognition situation during the speech recognition process, until the speech recognition is completed and the recognized text is output.
[0072] Understandably, before speech recognition begins, the speech recognition system will perform word frequency statistics on the input intervention hot word list and use the frequency of each hot word in the intervention hot word list as its initial recognition weight.
[0073] It is important to note that words in the first vocabulary are more closely related to the interview candidate's own information and are more likely to be mentioned in the interview answers. After determining the initial recognition weight of each hot word based on word frequency, the initial recognition weight of the hot words belonging to the first vocabulary can be further set, for example, by adding a fixed value to the already set initial recognition weight.
[0074] After determining the initial recognition weight for each hot word, the speech recognition system begins to perform speech recognition on the interview dialogue audio.
[0075] The recognition weight of hot words affects the decoding process, speech model, and acoustic model. During decoding, the decoder tends to select paths containing hot words with high recognition weights during the search process. In the language model, hot words with high recognition weights increase the probability of sentences or phrases containing these hot words, thus influencing the selection of decoding paths. When training the acoustic model, hot words with high recognition weights may cause the model to pay more attention to the acoustic features associated with these words.
[0076] Throughout the speech recognition process, the speech recognition system needs to dynamically adjust the recognition weight of each hot word in the intervention hot word list in real time based on the recognition of the interview dialogue audio. There are two ways to achieve this.
[0077] Method 1: For each hot word, if the speech recognition system detects that the hot word appears in the recognized text generated up to the current moment, then the recognition weight of the hot word is increased.
[0078] At each step of the speech recognition process, the speech recognition system monitors the generated recognized text and checks whether any hot words from the intervention hot word list appear in the generated recognized text. For hot words from the intervention hot word list that appear in the generated recognized text, the recognition weight of these hot words is increased.
[0079] When increasing the recognition weight of hot words, it can be achieved by following the preset weight increase rules, which can be either setting a weight increase multiple or adding a fixed value.
[0080] Method 2: For each hot word, if the speech recognition system detects that the phoneme of the hot word matches the segment of the interview dialogue audio being recognized at the current moment as well as a preset condition, then the recognition weight of the hot word is increased.
[0081] The speech recognition system segments the interview dialogue audio into phoneme-level segments and constructs a phoneme model for each hot word. Each hot word phoneme model contains the acoustic features of each phoneme in that hot word. During the speech recognition process, the system matches the phoneme segments of the interview dialogue audio with the phoneme models of each hot word in real time, calculating the matching degree (e.g., similarity score). For hot words with a matching degree exceeding a preset threshold, the recognition weight of these hot words is increased.
[0082] When increasing the weight of hot keywords, the extent of the weight increase can be determined based on the degree of relevance.
[0083] Understandably, speech recognition systems dynamically adjust the recognition weights of each hot word, making it easier for hot words with high recognition weights to be selected during the speech recognition process.
[0084] S103-2, the identified text is used as the converted text.
[0085] In this embodiment of the invention, the recognized text output by the speech recognition system can be directly used as the converted text corresponding to the interview dialogue audio.
[0086] To further improve the readability of the converted text of the interview dialogue audio, please refer to... Figure 3 The implementation process of step S103 may also include Figure 3 The sub-step S103-3 is shown.
[0087] S103-3 uses a pre-trained language model to perform error correction processing on the recognized text according to preset error correction instructions.
[0088] After obtaining the recognized text output by the speech recognition system, the recognized text and interview questions can be input into a pre-trained language model. The pre-trained language model performs semantic-level error correction on the recognized text according to preset error correction instructions, improving the accuracy and structure of the recognized text, and uses the corrected recognized text as the converted text corresponding to the interview dialogue audio.
[0089] Preset error correction instructions can include the following ordered instructions:
[0090] Instruction 1: If the original sentence already looks good, you don't need to change it.
[0091] Instruction 2: Retain the modal particles in the original sentence and keep their positions unchanged; maintain the sentence structure and word order unchanged.
[0092] Instruction 3: Determine if there are any misspellings in common phrases, such as translation errors due to differences in accent. If so, make the necessary corrections, prioritizing the use of appropriate words with similar pronunciations.
[0093] Instruction 4: Determine if there are any translation errors in technical terms. If so, correct them using the technical terms with the closest pronunciation in both Chinese and English.
[0094] Instruction 5: Everyone speaks differently in an interview, for example, regarding the issue of coherence. You need to assess whether the text contains appropriate punctuation. If the punctuation is severely inconsistent, you can add appropriate punctuation.
[0095] Instruction 6: Compare the text content before and after correction. If there are changes in the semantics of the content or significant changes in the candidate's speaking characteristics between the two versions, then abandon this correction instruction.
[0096] In order to perform the corresponding steps in the above method embodiments and various possible implementations, an implementation of an interview dialogue audio processing device 100 is given below.
[0097] Please refer to Figure 4 The interview dialogue audio processing device 100 may include an extraction module 101, a generation module 102, and a conversion module 103.
[0098] Extraction module 101 is used to extract first and second vocabulary from the resumes of interview candidates and the standard answers to interview questions using a pre-trained language model. The first and second vocabulary include common words and non-everyday words related to the position the interview candidate is applying for.
[0099] The generation module 102 is used to generate an intervention hot word list based on the first word and the second word.
[0100] The conversion module 103 is used to perform text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio.
[0101] Optionally, the generation module 102 is specifically used to add each word in the first vocabulary and the second vocabulary as a hot word to the intervention hot word list if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed a preset upper limit; and to add each word in the first vocabulary as a hot word to the intervention hot word list if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds a preset upper limit.
[0102] Optionally, the electronic device is equipped with a speech recognition system. The conversion module is specifically used to input the interview dialogue audio and the intervention hot word list into the speech recognition system, so that the speech recognition system can perform speech recognition on the interview dialogue audio according to each hot word in the intervention hot word list, and output the recognized text corresponding to the interview dialogue audio; the recognized text is used as the converted text.
[0103] Optionally, the conversion module 103 is further configured to use a pre-trained language model to perform error correction processing on the recognized text according to preset error correction instructions, and use the processed recognized text as the converted text.
[0104] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the interview dialogue audio processing device 100 described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0105] Furthermore, this embodiment of the invention also provides an electronic device 200, please refer to... Figure 5 The electronic device 200 may include a memory 210 and a processor 220.
[0106] The processor 220 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of a program for a method for processing interview dialogue audio provided in the above-described method embodiments.
[0107] The memory 210 may be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 210 may exist independently and be connected to the processor 220 via a communication bus. The memory 210 may also be integrated with the processor 220. The memory 210 is used to store machine-executable instructions for executing the scheme of this application. The processor 220 is used to execute the machine-executable instructions stored in the memory 210 to implement the above-described method embodiments.
[0108] This invention also provides a computer-readable storage medium containing a computer program, which, when executed, can be used to perform related operations in the interview dialogue audio processing method provided in the above-described method embodiments.
[0109] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for processing interview dialogue audio, characterized in that, Applied to electronic devices, wherein the electronic devices are equipped with a speech recognition system, the method includes: Using a pre-trained language model, the first and second vocabulary words are extracted from the resumes of interview candidates and the standard answers to interview questions, respectively. The first and second vocabulary words include common words and non-everyday words related to the job position applied for by the interview candidates. If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed the preset upper limit, then each word in the first vocabulary and the second vocabulary will be added as a hot word to the intervention hot word list. If the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds the preset upper limit, then each word in the first vocabulary will be added as a hot word to the intervention hot word list. Based on the aforementioned intervention hot word list, the interview dialogue audio is converted into text to obtain the converted text corresponding to the interview dialogue audio. The step of performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio includes: The interview dialogue audio and the intervention hot word list are input into the speech recognition system. Before performing speech recognition, the speech recognition system sets a recognition weight for each hot word according to its frequency of occurrence in the intervention hot word list. The speech recognition system performs speech recognition on the interview dialogue audio based on the recognition weight of each hot word, and dynamically adjusts the recognition weight of each hot word in real time according to the recognition situation during the speech recognition process, until the speech recognition is completed and the recognized text is output. The identified text is used as the converted text.
2. The method as described in claim 1, characterized in that, The steps of the speech recognition system to dynamically adjust the recognition weight of each hot word in real time based on the recognition results during the speech recognition process include: For each hot word, if the speech recognition system detects that the hot word appears in the recognized text generated up to the current time, the recognition weight of the hot word is increased.
3. The method as described in claim 1, characterized in that, The steps of the speech recognition system to dynamically adjust the recognition weight of each hot word in real time based on the recognition results during the speech recognition process include: For each of the hot words, if the speech recognition system detects that the phoneme of the hot word matches a segment of the interview dialogue audio being recognized at the current moment, the recognition weight of the hot word is increased.
4. The method as described in claim 1, characterized in that, Before using the identified text as the converted text, the step of performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio further includes: The pre-trained language model is used to perform error correction processing on the recognized text according to preset error correction instructions.
5. An audio processing device for interview dialogues, characterized in that, Applied to electronic devices, wherein the electronic devices are equipped with a voice recognition system, the device includes: The extraction module is used to extract the first vocabulary and the second vocabulary from the interview candidate's resume and the standard answer to the interview question, respectively, using a pre-trained language model. The first vocabulary and the second vocabulary include common words and non-everyday words related to the position the interview candidate is applying for. The generation module is configured to add each word in the first vocabulary and the second vocabulary as a hot word to the intervention hot word list if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary does not exceed a preset upper limit; and to add each word in the first vocabulary as a hot word to the intervention hot word list if the sum of the number of words in the first vocabulary and the number of words in the second vocabulary exceeds the preset upper limit. A conversion module is used to perform text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio. The step of performing text conversion processing on the interview dialogue audio based on the intervention hot word list to obtain the converted text corresponding to the interview dialogue audio includes: inputting the interview dialogue audio and the intervention hot word list into the speech recognition system, so that before performing speech recognition, the speech recognition system sets a recognition weight for each hot word according to the frequency of occurrence of each hot word in the intervention hot word list; the speech recognition system performs speech recognition on the interview dialogue audio according to the recognition weight of each hot word, and dynamically adjusts the recognition weight of each hot word in real time according to the recognition situation during the speech recognition process, until the speech recognition is completed and the recognized text is output; the recognized text is used as the converted text.
6. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the interview dialogue audio processing method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the interview dialogue audio processing method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Prompt text extension method and device, electronic equipment and storage medium
CN116738250A
Text error correction method and device, equipment and medium
CN117217207A
Speech recognition method and device, equipment and storage medium
CN118116388A