Text information display method and device, electronic equipment and storage medium
By displaying dialogue information in online training and converting users' voice data into text information, the problem of existing technologies being unable to simulate real communication scenarios is solved, thereby improving the authenticity and effectiveness of training.
Patent Information
- Application Number
- CN202211602551.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-09
AI Technical Summary
Existing online training methods cannot simulate real communication scenarios, resulting in poor training effectiveness.
By displaying dialogue information, the system receives the user's voice data and converts it into text information. The text display position is determined based on the position of the voice control, thus realizing the conversion and display of voice data into text information.
This improved the realism of the scenario-based exercises and enhanced the training effectiveness.
Smart Images

Figure CN116189682B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and more particularly, to a text information display method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the development of technology, people are increasingly learning online. For example, an enterprise can train employees over the network. However, the current training method over the network is too simple and cannot simulate real communication scenarios, resulting in poor training effects. SUMMARY
[0003] In view of the above problems, the present application provides a text information display method, device, electronic device, and storage medium to improve the above problems.
[0004] In a first aspect, the embodiments of the present application provide a text information display method, which includes: displaying conversation information; if voice data input by a user according to the conversation information is received, displaying a voice control corresponding to the voice data; determining a text display position according to position information of the voice control; converting the voice data into text information, and displaying the text information at the text display position.
[0005] In a second aspect, the embodiments of the present application also provide a text information display device, which includes: a first display module configured to display conversation information; a second display module configured to, if voice data input by a user according to the conversation information is received, display a voice control corresponding to the voice data; a determination module configured to determine a text display position according to position information of the voice control; and a conversion module configured to convert the voice data into text information, and display the text information at the text display position.
[0006] In a third aspect, the embodiments of the present application also provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text information display method of the first aspect.
[0007] In a fourth aspect, the embodiments of the present application also provide a computer-readable storage medium, which stores computer-executable instructions for enabling an electronic device to execute the text information display method of the first aspect.
[0008] The application provides a text information display method and device, electronic equipment and a storage medium. The method comprises the following steps: displaying conversation information; if voice data input by a user according to the conversation information is received, displaying a voice control corresponding to the voice data; determining a text display position according to position information of the voice control; converting the voice data into text information, and displaying the text information at the text display position. By automatically converting the voice data into text information and displaying the text information at the text display position, a real communication scenario can be simulated, the authenticity of scenario practice is improved, and the training effect is improved. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments of the present application, all other embodiments and drawings obtained by those skilled in the art without creative labor are within the scope of the present application.
[0010] Figure 1 is a flowchart of a text information display method provided by an embodiment of the present application.
[0011] Figure 2 is Figure 1 is a detailed flowchart of step 110 in
[0012] Figure 3 is a schematic diagram of a voice control and an input control provided by an embodiment of the present application.
[0013] Figure 4 is a schematic diagram of a display focus provided by an embodiment of the present application.
[0014] Figure 5 is another flowchart of a text information display method provided by an embodiment of the present application.
[0015] Figure 6 is a schematic diagram of an auxiliary control provided by an embodiment of the present application.
[0016] Figure 7 is a structural schematic diagram of a text information display device provided by an embodiment of the present application.
[0017] Figure 8 is a structural schematic diagram of an electronic equipment provided by an embodiment of the present application.
[0018] Figure 9 is a structural block diagram of a computer readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0020] With the development of technology, people are increasingly learning on the Internet. For example, enterprises can train employees through the network. However, the current training through the network is too simple. For example, many online training is just like answering questions to give users questions to input answers, which cannot simulate real communication scenarios. After the user completes the training, he or she still often encounters complaints when facing customers alone, the communication is not smooth, and the training effect is poor.
[0021] In order to improve the above problems, the inventors provide a text information display method, device, electronic equipment and storage medium. The method comprises: displaying conversation information; if voice data input by a user according to the conversation information is received, displaying a voice control corresponding to the voice data; determining a text display position according to position information of the voice control; converting the voice data into text information, and displaying the text information at the text display position. By automatically converting the voice data into text information and displaying the text information at the text display position, a real communication scenario can be simulated, the authenticity of scenario practice is improved, and the training effect is improved.
[0022] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0023] Please refer to Figure 1 , Figure 1 is a flowchart of a text information display method provided by an embodiment of the present application. As shown in Figure 1 , the method comprises steps 110 to 140.
[0024] Step 110: display conversation information.
[0025] In some embodiments, the conversation information is sent by a customer and displayed in the electronic device of the user.
[0026] In some embodiments, the conversation information can be pre-set, and the user replies to the conversation information by displaying the conversation information.
[0027] In some embodiments, the conversation information can be text information, audio information, video information and the like. The type of the conversation information is not limited in the present application.
[0028] Please refer to Figure 2 , Figure 2is Figure 1 a refinement process diagram of step 110 in FIG. 11. As shown in FIG. 11, step 110 includes steps 111-113. Figure 2
[0029] Step 111: Obtain scenario information and customer category information.
[0030] In some embodiments, in order to train the user, the scenario information and the customer category information can be obtained to determine the scenario for training and the target customer in the scenario.
[0031] In some embodiments, the scenario information includes scenario name, scenario attribute, etc. Optionally, the scenario attribute includes notice of payment, customer service, follow-up visit, product information update notice, etc.
[0032] Further, the scenario information can also include customer attribute. Exemplarily, the customer attribute includes whether to be in arrears, whether to be overdue, whether to have purchased, communication difficulty, etc.
[0033] In some embodiments, the customer category information is used to reflect the customer category, and further, one customer can correspond to multiple customer categories, so the customer category information can include multiple customer categories. Optionally, the customer category includes male, female, good credit, poor credit, domestic customer, foreign customer, etc.
[0034] In some embodiments, the historical scenario training information of the user can also be obtained, and one or more of the customer category, the customer attribute, and the scenario attribute which the user is least good at handling can be determined according to the historical scenario training information, so as to obtain the one or more of the customer category, the customer attribute, and the scenario attribute which the user is least good at handling, and to train the user in a targeted manner.
[0035] Step 112: Determine the target customer according to the scenario information and the customer category information.
[0036] In some embodiments, step 112 includes:
[0037] (1) Determine the customer attribute according to the scenario information.
[0038] (2) Determine the customer category according to the customer category information.
[0039] (3) Determine the target customer according to the customer attribute and the customer category.
[0040] In some embodiments, if multiple qualified customers can be determined according to the customer attribute and the customer category, one of the multiple qualified customers can be randomly selected as the target customer. Further, the dialogue information corresponding to each target customer is different.
[0041] In some embodiments, the administrator can set the scenario information in the background to set the customer attributes contained in the scenario information.
[0042] In some embodiments, the information of the target customer can be pre-created and stored, for example, can be pre-created and stored in a server, a database, or an electronic device of a user, and the target customer meeting the requirements can be found in the server, the database, or the electronic device of the user according to the customer attributes and the customer categories.
[0043] In some embodiments, in order to more accurately locate the target customer, the target customer can also be determined according to the scenario attributes, and at this time, step 112 comprises:
[0044] (1) determining the scenario attributes and the customer attributes according to the scenario information.
[0045] (2) determining the customer categories according to the customer category information.
[0046] (3) determining the target customer according to the scenario attributes, the customer attributes, and the customer categories.
[0047] Exemplarily, if the scenario attributes are the notice of payment reminder, the customer attributes are “whether overdue: yes”, and the customer categories are male and good credit, then the target customer is a male customer who has overdue and needs the notice of payment reminder and has good credit; if the scenario attributes are the notice of payment reminder, the customer attributes are “whether overdue: no”, and the customer categories are female and foreign customer, then the target customer is a foreign female customer who has not overdue and needs the notice of payment reminder.
[0048] According to the scenario attributes, the customer attributes, and the customer categories, a plurality of different target customers can be set, so that the target customer encountered by the user in each time of training is different, and the adaptability of the user is improved.
[0049] Step 113: determining the dialogue information according to the target customer.
[0050] In some embodiments, each target customer has corresponding dialogue information, and the dialogue information can be pre-set.
[0051] In some embodiments, the dialogue information is customized according to the customer attributes, the customer categories, and the scenario attributes of the target customer, so as to enhance the authenticity of the scenario practice.
[0052] Exemplarily, the dialogue information of the target customer with the customer attributes of “whether overdue: yes” can be the reason for overdue, and the dialogue information of the target customer with the customer attributes of “whether overdue: no” can be the impatience to the user for the information of payment reminder.
[0053] Exemplarily, the dialogue information of the target customer of the customer category of "domestic user" is all Chinese, and the dialogue information of the target customer of the customer category of "foreign user" is foreign language corresponding to the nationality of the foreign user or not very smooth Chinese, so as to train the communication ability of the user to different customers.
[0054] Exemplarily, the dialogue information is different when the scenario attribute is a notice of payment reminder, customer service, follow-up visit, and product information update notice, so as to simulate the real dialogue scenario.
[0055] In some embodiments, the user answers the displayed dialogue information by inputting voice data, and at this time, the displayed dialogue information is also related to the voice data input by the user last time.
[0056] Specifically, keyword detection can be performed on the voice data input by the user last time, or keyword detection can be performed on the text information obtained by voice recognition of the voice data input by the user last time, and the dialogue information to be displayed is determined according to the detected keywords and the target customer.
[0057] Specifically, the dialogue information is not fixed, in order to improve the authenticity of the scenario practice, when different keywords are detected, the dialogue information is also different, so as to simulate the real dialogue scene. For example, in the payment reminder scenario, when the detected keywords are "overdue" and "repayment time", the next displayed dialogue information is different. Exemplarily, when the keyword "overdue" is detected, the dialogue information can be to explain the reason for the overdue, and when the keyword "repayment time" is detected, the dialogue information is to explain whether the repayment can be made before the repayment time, so as to avoid the phenomenon that the displayed dialogue information is irrelevant to the voice data input by the user last time.
[0058] Further, emotion detection and / or timeout detection can also be performed on the voice data input by the user last time, and the dialogue information to be displayed is determined according to the detected keywords, emotions, whether the reply is in time, and the target customer. Exemplarily, when the target customer and the keyword are the same, when the user replies in time or has a good tone (for example, the detected emotion is gentle or the attitude is good), the displayed dialogue information is different from when the user replies overtime or has a bad tone (for example, the detected emotion is anger or sarcasm, etc.).
[0059] Exemplarily, the displayed dialogue information is difficult to reply when the user replies overtime or has a bad tone, for example, in the "payment reminder" scenario, the displayed dialogue information is "your attitude is too bad, I am very busy and have no time to continue communication with you!", and the displayed dialogue information is easy to reply when the user replies in time or has a good tone, for example, in the "payment reminder" scenario, the displayed dialogue information is "good, I will definitely repay the money within the specified time".
[0060] In some embodiments, the emotion recognition model can be used to detect the emotion of the voice data input by the user last time.
[0061] In some embodiments, a preset time threshold can be set. When the conversation information is displayed, if the user does not reply within the preset time threshold, for example, does not complete the input of voice data or trigger the input control within the preset time threshold, it is determined that the user reply is overdue. The input control is, for example, the input control 230 in Figure 3 , which will be described in the following part of the specification, and will not be repeated here. Further, the displayed conversation information can be adjusted according to the degree of user timeout. For example, the longer the timeout time is, the worse the attitude of the customer is.
[0062] Step 120: If the voice data input by the user according to the conversation information is received, display the voice control corresponding to the voice data.
[0063] In some embodiments, the user can input voice data through a microphone or other device.
[0064] In some embodiments, because the user can input multiple voice data, in order to improve the recognition accuracy of the voice data and prevent misrecognition of the voice data, one or more input controls are provided in the user's electronic device interface. When the user triggers the input control, it means that the user is inputting voice data according to the conversation information.
[0065] Please refer to Figure 3 , Figure 3 is a schematic diagram of the voice control and the input control provided by the embodiments of the present application. As shown in Figure 3 , the conversation interface 200 includes conversation information 210, voice control 220, and input control 230.
[0066] It can be understood that Figure 3 the position of the input control 230 in is only exemplary, and the present application does not limit the position of the input control 230.
[0067] In some embodiments, the position of the voice control 220 is below the conversation information 210. The present application does not limit the positional relationship between the voice control 220 and the conversation information 210.
[0068] In some embodiments, the user can trigger the voice control 220 to play the voice data input by the user.
[0069] In some embodiments, the style of the voice control 220 is pre-set in advance, and the voice control 220 is used to reflect that the user is inputting voice data when the voice control 220 is displayed.
[0070] In some embodiments, after the user triggers the input control 230, it is indicated that the user is inputting voice data according to the dialogue information 210, and at this time, the voice control 220 is displayed.
[0071] Step 130: determining the text display position according to the position information of the voice control.
[0072] In some embodiments, the position information of the display focus can be determined according to the position information of the voice control, and the text display position is determined according to the position information of the display focus. Optionally, the position of the display focus has a first correspondence relationship with the position of the voice control, and when the position information of the voice control is determined, the position information of the display focus can be determined according to the first correspondence relationship.
[0073] In some embodiments, after the position information of the display focus is determined, the position of the display focus can be directly taken as the text display position.
[0074] In some embodiments, please refer to Figure 4 , Figure 4 is a schematic diagram of the display focus provided by the embodiments of the present application. As shown in Figure 4 , Figure 4 comprises dialogue information 210, voice control 220, input control 230 and display focus 240, wherein the dialogue information 210, voice control 220 and input control 230 have been introduced in the rest of the description, and will not be repeated here.
[0075] In some embodiments, the display focus 240 has a preset style, for example Figure 4 , so as to prompt the user that the voice data will be displayed where after being converted into text information.
[0076] Step 140: converting the voice data into text information, and displaying the text information at the text display position.
[0077] In some embodiments, the final display effect of the text display position will change according to the information amount of the text information, for example, when the text contains more words, the area of the text display position is larger; when the text contains fewer words, the area of the text display position is smaller.
[0078] In some embodiments, the voice data can be recognized according to a preset trained voice recognition model to obtain the text information.
[0079] In some embodiments, after step 140, the text information display method provided by the embodiments of the present application further comprises:
[0080] (1) analyzing the voice data to obtain voice analysis results.
[0081] (2) determining a user's answer score according to the voice analysis result; wherein the answer score is used to reflect whether the user's answer meets the requirements.
[0082] In some embodiments, the voice data can be analyzed according to the trained voice analysis model to obtain the voice analysis result.
[0083] In some embodiments, the voice analysis includes one or more of speech rate detection, volume detection, emotion recognition, and whether the reply is in time.
[0084] Exemplarily, the voice analysis result corresponding to the speech rate detection can be too low speech rate, normal speech rate, or too high speech rate; or a specific value of the speech rate, the unit of which can be words per minute.
[0085] Further, a baseline score can be preset for the user, and the baseline score is deducted based on the preset speech rate range to obtain the user's answer score when the user's specific speech rate is outside the preset speech rate range; or the baseline score is deducted based on the user's specific speech rate when the user's specific speech rate is outside the preset speech rate range, for example, the further the user's specific speech rate is from the preset speech rate range, the more the score is deducted.
[0086] Exemplarily, the voice analysis result corresponding to the volume detection can be too low volume, normal volume, or too high volume; or a specific value of the volume, the unit of which can be decibels.
[0087] Further, a baseline score can be preset for the user, and the baseline score is deducted based on the preset volume range to obtain the user's answer score when the user's specific volume is outside the preset volume range; or the baseline score is deducted based on the user's specific volume when the user's specific volume is outside the preset volume range, for example, the further the user's specific volume is from the preset volume range, the more the score is deducted.
[0088] Exemplarily, the voice analysis result corresponding to the emotion recognition can be abnormal emotion or normal emotion; or a specific emotion, for example, the specific emotion is recognized as anger, gentleness, or impatience.
[0089] Further, a baseline score can be preset for the user, and the baseline score is deducted based on the preset volume range to obtain the user's answer score when the user's specific volume is outside the preset volume range; or the baseline score is deducted based on the user's specific volume when the user's specific volume is outside the preset volume range, for example, the further the user's specific volume is from the preset volume range, the more the score is deducted.
[0090] Exemplarily, the voice analysis result corresponding to whether the reply is in time can be a reply in time or a reply not in time; or a specific timeout time, the unit of which can be seconds.
[0091] Further, a reference score can be preset for the user, and when the user replies beyond the time limit, the reference score is deducted based on the reference score, and when the user replies within the time limit, no deduction is made, so as to obtain the answer score of the user; or the time of the user's reply beyond the time limit is used to deduct the score, and the longer the time of the user's reply beyond the time limit, the more the score is deducted.
[0092] In some embodiments, the method for displaying text information further comprises:
[0093] (1) analyzing the text information to obtain a text analysis result.
[0094] (2) determining the answer score of the user according to the voice analysis result and the text analysis result.
[0095] In some embodiments, the text information can be analyzed according to the trained text information analysis model to obtain the text analysis result.
[0096] In some embodiments, the text information analysis includes one or more of keyword detection, sensitive word detection, and dialogue accuracy.
[0097] For example, the text analysis result corresponding to the keyword detection can be whether the text information contains the keyword or not.
[0098] Further, a reference score can be preset for the user, and when the text information does not contain the keyword, the reference score is deducted based on the reference score, and the deduction of the voice analysis result is superimposed, so as to obtain the answer score of the user.
[0099] For example, the text analysis result corresponding to the sensitive word detection can be whether the text information contains the sensitive word or not; or can be a specific sensitive word.
[0100] Further, a reference score can be preset for the user, and when the text information contains the sensitive word, the reference score is deducted based on the reference score, and the deduction of the voice analysis result is superimposed, so as to obtain the answer score of the user; or according to the deduction value corresponding to each sensitive word, the reference score is deducted based on the reference score, and the deduction of the voice analysis result is superimposed, so as to obtain the answer score of the user, wherein the deduction value corresponding to each sensitive word can be different.
[0101] In some embodiments, after the sensitive word detection, the sensitive word can be highlighted in the text information to inform the user why the sensitive word is deducted. Further, different sensitive words with different deduction values can be highlighted in different ways, for example, the sensitive word with the highest deduction value is highlighted in red, and the sensitive word with the second highest deduction value is highlighted in yellow. It can be understood that the sensitive word can also be highlighted in other ways such as bold or italic, and the specific way of highlighting the sensitive word is not limited in the present application.
[0102] Exemplarily, the text analysis result corresponding to the dialogue accuracy can be a similarity of the text information and the reference text, for example, can be a similarity percentage of the text information and the reference text, for example, the similarity of the text information and the reference text is 80%, and the reference text will be introduced in the following part of the specification.
[0103] Further, a reference score can be preset for the user, when the similarity percentage of the text information and the reference text reaches the preset percentage, no deduction is made, when the similarity percentage of the text information and the reference text does not reach the preset percentage, deduction is made, and the greater the difference between the similarity percentage and the preset percentage, the more deduction is made, and the deduction of the voice analysis result is superimposed to obtain the answer score of the user.
[0104] In the above manner, the answer score of the user can be obtained to visualize the training effect of the user, and the answer score can be obtained according to the specific voice analysis and text analysis result, so that the user can optimize the deduction item again when the scenario training is performed again, and the training effect of the scenario training is improved.
[0105] In some embodiments, please refer to Figure 5 , Figure 5 is another flowchart of a text information display method provided by an embodiment of the present application. As shown in Figure 5 , the text information display method 100 includes steps 110 to 180.
[0106] Step 110: display dialogue information.
[0107] Specifically, step 110 has been introduced in the rest of the specification, and will not be described here.
[0108] Step 120: if voice data input by the user according to the dialogue information is received, display a voice control corresponding to the voice data.
[0109] Specifically, step 120 has been introduced in the rest of the specification, and will not be described here.
[0110] Step 130: determine a text display position according to position information of the voice control.
[0111] Specifically, step 130 has been introduced in the rest of the specification, and will not be described here.
[0112] Step 140: convert the voice data into text information, and display the text information at the text display position.
[0113] Specifically, step 140 has been introduced in the rest of the specification, and will not be described here.
[0114] Step 150: display the auxiliary control when receiving the voice data input by the user according to the dialogue information.
[0115] In some embodiments, please refer to Figure 6 , Figure 6 is a schematic diagram of the auxiliary control provided by the embodiments of the present application. As shown in Figure 6 ,the auxiliary control 250 includes dialogue information 210, voice control 220, input control 230, display focus 240, and the auxiliary control 250, wherein the dialogue information 210, voice control 220, input control 230, and display focus 240 have been described in the rest of the specification and will not be repeated here. Figure 6
[0116] In some embodiments, the style of the auxiliary control can be set as needed.
[0117] In some embodiments, the position of the auxiliary control is determined according to the display focus 240 and the dialogue information 210, for example, the position of the auxiliary control is located between the display focus 240 and the dialogue information 210, and the auxiliary control does not overlap with the display focus 240 and the dialogue information 210, etc.
[0118] Step 160: determine the reference text display position according to the position information of the auxiliary control.
[0119] In some embodiments, after determining the position information of the auxiliary control, the position of the auxiliary control can be directly used as the reference text display position.
[0120] Step 170: determine the reference text according to the dialogue information.
[0121] In some embodiments, different dialogue information corresponds to different reference texts.
[0122] In some embodiments, the second correspondence relationship between the dialogue information and the reference text can be set in advance, and the reference text corresponding to the dialogue information can be determined according to the second correspondence relationship after displaying the dialogue information.
[0123] Further, the reference text can also contain keywords, when the voice data input by the user or the text information obtained according to the voice data contains the keywords in the reference text, the next dialogue information will be displayed after the user inputs the voice data, so as to prevent the user from answering questions randomly or invalidly.
[0124] In some embodiments, different levels can also be set for different keywords, and different styles can be used to display keywords of different levels, for example, different colors, different fonts, or different ways such as bold, italic, etc. to distinguish keywords of different levels.
[0125] In some embodiments, the keywords related to the business can be set as the highest level, and the keywords related to the courtesy can be set as the lower level.
[0126] Step 180: displaying the reference text at the reference text display position.
[0127] In some embodiments, the final display effect of the reference text display position can vary according to the information amount of the reference text information, for example, when the reference text contains more words, the area of the reference text display position is larger; when the reference text contains fewer words, the area of the reference text display position is smaller.
[0128] In the above manner, the reference text can be displayed at the reference text display position to provide visual reference text for the user during training, helping the user quickly master the standard dialogue language and improving the training effect of the scene practice.
[0129] In some embodiments, the text information display method provided by the embodiments of the present application further includes:
[0130] (1) When the voice data input by the user according to the dialogue information is received, performing voiceprint recognition on the voice data to obtain the first voiceprint feature corresponding to the voice data.
[0131] (2) Obtaining the second voiceprint feature corresponding to the user.
[0132] (3) Performing identity verification on the user according to the first voiceprint feature and the second voiceprint feature, and saving the identity verification result.
[0133] In some embodiments, the voiceprint recognition model is trained to perform voiceprint recognition on the voice data to obtain the first voiceprint feature corresponding to the voice data.
[0134] In some embodiments, the user also needs to perform voice input when registering an account, and the second voiceprint feature can be obtained by performing voiceprint recognition on the voice input by the user during registration; wherein the second voiceprint feature can be saved in the server or the database, so as to obtain the second voiceprint feature in the server or the database after the user inputs the voice data.
[0135] In some embodiments, if the first voiceprint feature matches the second voiceprint feature, for example, the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset similarity, such as the similarity between the first voiceprint feature and the second voiceprint feature is greater than 80%, it means that the user is the person who is answering, and the identity verification is successful; if the first voiceprint feature does not match the second voiceprint feature, it means that someone is taking the place of the user to answer, that is, the user is not the person who is answering, and the identity verification fails.
[0136] In some embodiments, after saving the authentication result, the administrator can view the authentication result in the background to inquire or punish the user who fails in the authentication.
[0137] In this way, the remaining people can be prevented from taking the place of the user to perform the scenario practice, and the difficulty of cheating is improved.
[0138] The application provides a text information display method, which comprises the following steps: displaying conversation information; if voice data input by a user according to the conversation information is received, displaying a voice control corresponding to the voice data; determining a text display position according to position information of the voice control; converting the voice data into text information, and displaying the text information at the text display position. By automatically converting the voice data into text information and displaying the text information at the text display position, a real communication scenario can be simulated, the authenticity of the scenario practice is improved, and the training effect is improved.
[0139] Please refer to Figure 7 , Figure 7 is a structural schematic diagram of a text information display device provided by an embodiment of the application. As shown in Figure 7 , the text information display device 300 comprises a first display module 310, a second display module 320, a determination module 330 and a conversion module 340.
[0140] The first display module 310 is configured to display conversation information.
[0141] The second display module 320 is configured to, if voice data input by a user according to the conversation information is received, display a voice control corresponding to the voice data.
[0142] The determination module 330 is configured to determine a text display position according to position information of the voice control.
[0143] The conversion module 340 is configured to convert the voice data into text information, and display the text information at the text display position.
[0144] It should be noted that, for the device embodiment, the description is relatively simple because the device embodiment is basically similar to the method embodiment, and the relevant parts can be referred to the part of the description of the method embodiment. For any processing mode described in the method embodiment, the corresponding processing module can be used in the device embodiment, and the device embodiment will not be described again.
[0145] In addition, each functional module in each embodiment of the application can be integrated in one processing module, or each module can exist physically, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.
[0146] Please refer to Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 400 includes: one or more processors 410 and a memory 420. Figure 8 Take a processor 410 as an example.
[0147] The processor 410 and the memory 420 can be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.
[0148] The processor 410 is used to display dialogue information; if it receives voice data input by the user based on the dialogue information, it displays the voice control corresponding to the voice data; it determines the text display position based on the position information of the voice control; it converts the voice data into text information and displays the text information at the text display position.
[0149] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules of the text information display method in the embodiments of this application. The processor 410 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 420, thereby implementing the text information display method of the above-described method embodiments.
[0150] The memory 420 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include memory remotely located relative to the processor 410, and these remote memories may be connected to the controller via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0151] One or more modules are stored in memory 420. When executed by one or more processors 410, they perform the text information display method in any of the above method embodiments, for example, the method described above. Figure 1 Steps 110 to 140 of the method.
[0152] Please refer to Figure 9 , Figure 9is a structural block diagram of a computer readable storage medium provided by an embodiment of the present application. The computer readable storage medium 500 stores program code 510, which can be invoked by a processor to execute the text information display method described in the above method embodiments.
[0153] The computer readable storage medium 500 can be an electronic storage such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk or a ROM. Alternatively, the computer readable storage medium includes a non-transitory computer readable medium. The computer readable storage medium 500 has a storage space for program code for executing any method step of the above text information display method. These program codes can be read from or written into one or more computer program products. The program codes can be compressed in a suitable form, for example.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; under the idea of the present application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the present application as described above. In order to be brief, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application. Through the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course, they can also be implemented by hardware. Those skilled in the art can understand that all or part of the processes in the above embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium, and when the program is executed, it can include the processes of the above embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
Claims
1. A text information display method characterized by comprising: The method includes: Display dialogue information, which is pre-set information; If voice data is received from the user based on the dialogue information, a voice control corresponding to the voice data is displayed. The displayed voice control is used to reflect that voice data is being input. The text display position is determined based on the position information of the voice control; The voice data is automatically converted into text information and displayed at the text display location; When the text information includes keywords from the reference text corresponding to the dialogue information, the next dialogue information is displayed.
2. The method of claim 1, wherein, Before displaying the dialogue information, the method further includes: Obtain contextual information and customer category information; Target customers are determined based on the context information and the customer category information; Determine the dialogue information based on the target customer.
3. The method of claim 2, wherein, The step of determining the target customer based on the context information and the customer category information includes: Determine customer attributes based on the aforementioned scenario information; Determine the customer category based on the customer category information; The target customer is determined based on the customer attributes and the customer category.
4. The method of claim 1, wherein, The method further includes: When the voice data input by the user based on the dialogue information is received, the auxiliary control is displayed; The reference text display position is determined based on the position information of the auxiliary control; The reference text is determined based on the dialogue information; The reference text is displayed at the specified reference text display location.
5. The method of claim 1, wherein, After displaying the text information at the designated text display location, the method further includes: The speech data is analyzed to obtain speech analysis results; The user's response score is determined based on the voice analysis results; wherein the response score is used to reflect whether the user's response meets the requirements.
6. The method of claim 5, wherein, Determining the user's response score based on the voice analysis results includes: The text information is analyzed to obtain the text analysis results; The user's answer score is determined based on the voice analysis results and the text analysis results.
7. The method of claim 1, wherein, The method further includes: Upon receiving voice data input by the user based on the dialogue information, voiceprint recognition is performed on the voice data to obtain the first voiceprint feature corresponding to the voice data; Obtain the second voiceprint feature corresponding to the user; The user is authenticated based on the first voiceprint feature and the second voiceprint feature, and the authentication result is saved.
8. A text information display device, characterized by comprising: The device includes: The first display module is used to display dialogue information, which is pre-set information. The second display module is used to display a voice control corresponding to the voice data if it receives voice data input by the user based on the dialogue information. The displayed voice control is used to reflect that voice data is being input. The determination module is used to determine the text display position based on the position information of the voice control; The conversion module is used to automatically convert the voice data into text information and display the text information at the text display position; The device is also used to display the next dialogue information when the text information includes keywords from the reference text corresponding to the dialogue information.
9. An electronic device, comprising: include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the text information display method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that enable the electronic device to perform the text information display method according to any one of claims 1-7.
Citation Information
Patent Citations
Training method and device, computer equipment and storage medium
CN110458732A
Interaction method and electronic equipment
CN111147444A
Dialogue training method and device based on voice performance portrait, and electronic equipment
CN115309917A