Voice information processing method and device, electronic equipment and storage medium

By using a voice information processing system between bank staff and customers to identify, encrypt, and convert the voice information of the target object into text information, the problems of low efficiency and poor security of information interaction between bank customers and staff are solved, and the efficiency and security of information interaction through various interaction methods are improved.

CN116645965BActive Publication Date: 2025-12-16INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310629036.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-30
Publication Date
2025-12-16
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

In existing technologies, when bank staff and bank customers interact via voice through audio devices, the information exchange efficiency is low, and there are problems with information leakage and poor security.

Method used

The system collects voice information from multiple objects using audio acquisition devices, identifies the voice information of the target object and converts it into text information, identifies and encrypts the information to be encrypted, and sends it to the target projection device for display. Combining voice feature recognition and keyword detection, it provides multiple interaction methods to improve efficiency and security.

Benefits of technology

It improves the efficiency of information interaction, avoids the problems of low efficiency and poor security, provides text information interaction methods in addition to voice interaction, and enhances information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645965B_ABST
    Figure CN116645965B_ABST
Patent Text Reader

Abstract

The application discloses a voice information processing method and device, electronic equipment and storage medium, and relates to the field of financial technology and other related technical fields. The method comprises the following steps: collecting voice information of N objects through an audio acquisition device, and identifying voice information of a target object from the voice information of the N objects; converting the voice information of the target object into first text information; identifying whether there is to-be-encrypted information in the first text information; in the case where there is to-be-encrypted information in the first text information, performing a data processing operation on the first text information to obtain second text information, wherein the data processing operation is used for encrypting the to-be-encrypted information in the first text information; and sending the second text information to a target screen projection device for display. The application solves the technical problem of low information interaction efficiency caused by only using an audio device for information interaction in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of financial technology and other related technical fields, in particular, to a voice information processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] In the existing bank working scene, the glass on the counter ensures the safety of the bank staff and the bank customers, but also reduces the information interaction efficiency between the bank staff and the bank customers. For example, in the prior art, the bank staff and the bank customers usually interact through the microphone and the audio playing device on the counter, but this information interaction mode is prone to the problem that the bank customer cannot hear the staff's speech or the staff cannot hear the bank customer's speech, thereby causing low information interaction efficiency between the two parties.

[0003] For the above problems, no effective solution has been proposed so far. SUMMARY

[0004] The present application provides a voice information processing method and device, electronic equipment and storage medium to at least solve the technical problem of low information interaction efficiency caused by only using audio devices for information interaction in the prior art.

[0005] According to one aspect of the present application, a voice information processing method is provided, comprising: collecting voice information of N objects through an audio collection device, and identifying voice information of a target object from the voice information of the N objects, wherein N is a positive integer, and the target object is an object corresponding to a target screen projection device; converting the voice information of the target object into first text information; identifying whether there is to-be-encrypted information in the first text information, wherein the to-be-encrypted information is information that is prohibited from being publicly displayed through the target screen projection device; in the case where there is to-be-encrypted information in the first text information, performing a data processing operation on the first text information to obtain second text information, wherein the data processing operation is used to encrypt the to-be-encrypted information in the first text information; and sending the second text information to the target screen projection device for display.

[0006] Further, the voice information processing method further comprises: performing sound feature recognition processing on the voice information of the N objects through a target model to obtain the sound feature of each object, wherein the target model is a neural network model trained using voice information of M objects with known sound features as training samples, and the M objects at least include the N objects; and taking the voice information corresponding to the sound feature of the target object as the voice information of the target object.

[0007] Further, the method further includes: identifying first voice information in the voice information of the target object, the first voice information being voice information with a volume less than a preset volume; removing the first voice information from the voice information of the target object to obtain second voice information corresponding to the target object; and converting the second voice information into first text information.

[0008] Further, the method further includes: identifying whether a first keyword exists in the first text information, the first keyword being used to represent information that is prohibited from being displayed to the first object, the first object being an object viewing the target screen projection device; in a case where the first keyword exists in the first text information, determining that information corresponding to the first keyword in the first text information is to be encrypted information; and in a case where the first keyword does not exist in the first text information, determining that the to-be-encrypted information does not exist in the first text information.

[0009] Further, the method further includes: after the voice information of the target object is converted into the first text information, detecting whether a second keyword exists in the first text information, the second keyword being a keyword related to a business form, the business form being a form that needs to be filled out or reviewed by the first object; in a case where the second keyword exists in the first text information, sending the business form to the target screen projection device for display; and in a case where the second keyword does not exist in the first text information, prohibiting the business form from being sent to the target screen projection device for display.

[0010] Further, the method further includes: after the voice information of the target object is converted into the first text information, detecting whether a third keyword exists in the first text information, the third keyword being a keyword related to a target financial product, the target financial product being a financial product recommended by the target object to the first object; in a case where the third keyword exists in the first text information, sending product information of the target financial product to the target screen projection device for display; and in a case where the third keyword does not exist in the first text information, prohibiting the product information of the target financial product from being sent to the target screen projection device for display.

[0011] Further, the method further includes: after sending the second text information to the target projection device for display, obtaining voice information of the first object; identifying whether a fourth keyword exists in the voice information of the first object, the fourth keyword being used to represent a consultation question raised by the first object to the target object with respect to the second text information; in a case where the fourth keyword exists in the voice information of the first object, performing a differentiated text operation on target text information in the second text information, the target text information being text information related to the fourth keyword, the differentiated text operation being used to switch a text format of the target text information to a text format different from a text format of other text information, the other text information being text information in the second text information other than the target text information; and in a case where the fourth keyword does not exist in the voice information of the first object, prohibiting the differentiated text operation on the target text information in the second text information.

[0012] According to another aspect of the present application, a voice information processing apparatus is also provided, which includes: a collection module configured to collect voice information of N objects through an audio collection device, and identify voice information of a target object from the voice information of the N objects, N being a positive integer, and the target object being an object corresponding to a target projection device; a conversion module configured to convert the voice information of the target object into first text information; an identification module configured to identify whether to-be-encrypted information exists in the first text information, the to-be-encrypted information being information that is prohibited from being publicly displayed through the target projection device; a data processing module configured to, in a case where the to-be-encrypted information exists in the first text information, perform a data processing operation on the first text information to obtain second text information, the data processing operation being used to encrypt the to-be-encrypted information in the first text information; and an information sending module configured to send the second text information to the target projection device for display.

[0013] According to another aspect of the present application, a computer readable storage medium having a computer program stored therein is also provided, wherein the computer program, when executed, controls a device in which the computer readable storage medium is located to perform the voice information processing method described above.

[0014] According to another aspect of the present application, an electronic device is also provided, which includes one or more processors and a memory, the memory being configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement the voice information processing method described above.

[0015] In the present application, the voice information of the target object is converted into text information, the voice information of N objects is collected by an audio collection device, the voice information of the target object is recognized from the voice information of the N objects, the voice information of the target object is then converted into first text information, it is then determined whether there is to-be-encrypted information in the first text information, in the case where there is to-be-encrypted information in the first text information, data processing operation is performed on the first text information to obtain second text information, and finally the second text information is sent to the target screen projection device for display. Wherein, N is a positive integer, the target object is an object corresponding to the target screen projection device; the to-be-encrypted information is information that is prohibited from being publicly displayed on the target screen projection device; the data processing operation is used for encrypting the to-be-encrypted information in the first text information.

[0016] From the above, the present application converts the voice information of the target object into text information and displays it on the target screen projection device corresponding to the target object, thereby providing a text information interaction mode in addition to voice interaction, and thus the technical problem of low information interaction efficiency in the prior art that only uses an audio device for information interaction can be avoided. In addition, the present application also recognizes the voice information of the target object from the voice information of N objects, thereby avoiding displaying the voice information of other bank staff who do not correspond to the target screen projection device on the target screen projection device. Furthermore, the present application can also determine whether there is to-be-encrypted information in the first text information, and in the case where there is to-be-encrypted information in the first text information, data encryption processing is performed on the first text information, thereby avoiding the problem of poor information security caused by indiscriminately converting the voice information of the target object into text information.

[0017] Therefore, the technical scheme of the present application achieves the purpose of information interaction through multiple interaction modes, thereby achieving the technical effect of improving information interaction efficiency, and thus solving the technical problem of low information interaction efficiency in the prior art that only uses an audio device for information interaction. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0019] Figure 1 is a flowchart of an optional voice information processing method according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of an optional conversion of voice information of a target object into first text information according to an embodiment of the present application;

[0021] Figure 3is a schematic diagram of an optional first device according to an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of an optional second device according to an embodiment of the present application;

[0023] Figure 5 is a flow chart of another voice information processing method according to an embodiment of the present application;

[0024] Figure 6 is a schematic diagram of an optional voice information processing apparatus according to an embodiment of the present application;

[0025] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to the process, method, product or device.

[0028] It should also be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties. For example, an interface is provided between the system and the relevant user or institution. Before obtaining the relevant information, the interface needs to send a request for obtaining the relevant information to the aforementioned user or institution, and after receiving the consent information fed back by the aforementioned user or institution, the relevant information is obtained.

[0029] The present application will be further illustrated below in conjunction with various embodiments.

[0030] Embodiment 1

[0031] According to an embodiment of the present application, an embodiment of a method for processing voice information is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] Figure 1 is a flowchart of an optional method for processing voice information according to an embodiment of the present application, as shown in Figure 1 The method comprises the following steps:

[0033] In step S101, voice information of N objects is collected by an audio collection device, and voice information of a target object is identified from the voice information of the N objects.

[0034] In step S101, N is a positive integer, and the target object is an object corresponding to a target screen projection device.

[0035] Optionally, the method for processing voice information in the present application can be applied in the scenario of a bank branch performing business for a bank user. In the existing bank branch, a plurality of counters are usually provided, wherein a glass is arranged on each counter, and the two sides of the glass are seats arranged for bank staff and bank users respectively, so that the bank staff and the bank user can communicate across the glass. Such a setting is to ensure the personal safety and financial safety of the bank staff and the bank user.

[0036] Further, in the prior art, in order to ensure that the bank staff and the bank user on both sides of the glass can communicate information, an audio device is further arranged on the counter, which is used to collect the voice of the bank staff and play it to the bank user, and collect the voice of the bank user and play it to the bank staff. On this basis, the present application provides a voice information processing system for executing the method for processing voice information in the present application, wherein the voice information processing system can run in the audio device on the counter in the form of a software system or an embedded system.

[0037] It should be noted that each service window on the counter is provided with an audio device and a screen projection device, and the bank staff using the service window has a corresponding relationship with the audio device and the screen projection device of the service window. For example, the counter of a certain bank outlet has 4 windows, and the bank staff responsible for the 4 windows are object 1, object 2, object 3 and object 4, wherein object 1 is responsible for window A and corresponds to audio device A-1 and screen projection device A-2 of window A, object 2 is responsible for window B and corresponds to audio device B-1 and screen projection device B-2 of window B, object 3 is responsible for window C and corresponds to audio device C-1 and screen projection device C-2 of window C, and object 4 is responsible for window D and corresponds to audio device D-1 and screen projection device D-2 of window D. Therefore, in this example, the N objects include at least object 1, object 2, object 3 and object 4.

[0038] Further, the audio acquisition device can be an audio acquisition device in the audio device of any one window, for example, for window A, the audio acquisition device corresponding to window A can be microphone A-1-1 in audio device A-1 for collecting the voice of the bank staff, and object 1 is the target object corresponding to microphone A-1-1, and object 1 also corresponds to screen projection device A-2 of window A.

[0039] It is easy to understand that when a bank customer comes to window A to handle business, he only needs to communicate with object 1, so when microphone A-1-1 receives the voice information of multiple objects, the voice information of multiple objects except the voice information of object 1 is redundant voice information, which is not needed by the bank customer of window A. In addition, the voice information of other objects may also involve private data, for example, the voice information of object 2 may contain private data of the bank customer of window B, so if the voice information of object 2 is conveyed to the bank user of window A, it may also cause the problem of leaking the private data of the bank customer of window B. On this basis, the present application identifies the voice information of the target object from the voice information of the N objects, since the target object is the object corresponding to the target screen projection device, so it can not only avoid that the bank customer receives redundant voice information, but also improve the information security of the bank customer.

[0040] In step S102, the voice information of the target object is converted into first text information.

[0041] Optionally, in step S102, after the voice information of the target object is identified, the voice information processing system can perform voice-to-text processing on the voice information of the target object to obtain text information (first text information) corresponding to the voice information of the target object.

[0042] Step S103, identifying whether the to-be-encrypted information exists in the first text information.

[0043] In step S103, the to-be-encrypted information is information that is prohibited from being publicly displayed through the target screen projection device.

[0044] Optionally, in some scenarios, there can also be some private data in the voice information of the target object. For example, when the target object is the object 1, the object 1 can need the assistance of other objects when handling some businesses. For example, the object 1 needs the authority confirmation of the person in charge of the bank outlet, and therefore, in the communication process between the object 1 and the other objects, some information that is not allowed to be publicly disclosed within the bank can be involved. If this part of information is also directly displayed on the target screen projection device, it can cause the risk of information leakage. In addition, the to-be-encrypted information can also be private data of a bank customer. For example, when the object 1 communicates with the bank customer of the window A, the object 1 can confirm the identity card number and other private data of the bank customer through voice. Since this information can be seen by other bank customers when being displayed on the screen, the information can also be encrypted and then displayed as to-be-encrypted information.

[0045] Step S104, in the case where the to-be-encrypted information exists in the first text information, performing a data processing operation on the first text information to obtain second text information.

[0046] In step S104, the data processing operation is used to encrypt the to-be-encrypted information in the first text information.

[0047] Optionally, after identifying that the to-be-encrypted information exists in the first text information, the voice information processing system can encrypt the to-be-encrypted information through a preset encryption algorithm. For example, the to-be-encrypted information can be uniformly converted into special characters.

[0048] Step S105, sending the second text information to the target screen projection device for display.

[0049] Optionally, after generating the second text information, the voice information processing system can send the second text information to the target screen projection device for display, so as to be viewed by the bank customer. It should be noted that the target screen projection device can be a display screen.

[0050] Based on the contents of the above steps S101 to S105, in the present application, the voice information of the target object is converted into text information, the voice information of N objects is collected by an audio collection device, the voice information of the target object is recognized from the voice information of the N objects, the voice information of the target object is converted into first text information, it is then determined whether there is to-be-encrypted information in the first text information, in the case that there is to-be-encrypted information in the first text information, the first text information is subjected to data processing operation to obtain second text information, and finally the second text information is sent to the target screen projection device for display. Wherein, N is a positive integer, the target object is an object corresponding to the target screen projection device; the to-be-encrypted information is information that is prohibited from being publicly displayed on the target screen projection device; the data processing operation is used for encrypting the to-be-encrypted information in the first text information.

[0051] From the above content, the present application converts the voice information of the target object into text information and displays it on the target screen projection device corresponding to the target object, thereby providing a text information interaction mode in addition to voice interaction, thereby avoiding the technical problem of low information interaction efficiency in the prior art that only audio devices are used for information interaction. In addition, the present application also recognizes the voice information of the target object from the voice information of the N objects, thereby avoiding displaying the voice information of other bank staff who are not corresponding to the target screen projection device on the target screen projection device. In addition, the present application can also recognize whether there is to-be-encrypted information in the first text information, and in the case that there is to-be-encrypted information in the first text information, the first text information is subjected to data encryption processing, thereby avoiding the problem of poor information security caused by indiscriminately converting the voice information of the target object into text information.

[0052] Therefore, the technical scheme of the present application achieves the purpose of information interaction through multiple interaction modes, thereby achieving the technical effect of improving the information interaction efficiency, and further solving the technical problem of low information interaction efficiency caused by only using audio devices for information interaction in the prior art.

[0053] In an optional embodiment, the voice information processing system can perform sound feature recognition processing on the voice information of the N objects by a target model to obtain the sound feature of each object, wherein the target model is a neural network model trained by using the voice information of M objects with known sound features as training samples, and the M objects at least include the N objects. Then, the voice information processing system takes the voice information corresponding to the sound feature of the target object as the voice information of the target object.

[0054] Optionally, to obtain the target model, voice information from M objects (e.g., all employees) at a bank branch can be pre-collected as training samples. Then, the voice features of each object are extracted based on its voice information, resulting in voice features for each object. These voice features are used as label data corresponding to the object's voice information. Subsequently, the voice information of each object and its corresponding label data are input into a pre-set deep learning neural network and trained iteratively multiple times to obtain the target model. The target model can output the voice features corresponding to each input voice information.

[0055] Optionally, the target model is deployed in the speech information processing system, and the speech information processing system, knowing the target object corresponding to the current window, will use the speech information corresponding to the speech features of the target object as the speech information of the target object based on the speech features of each object output by the target model.

[0056] In one alternative embodiment, Figure 2 A flowchart illustrating an optional method for converting speech information of a target object into first text information according to an embodiment of this application is shown, including the following steps:

[0057] Step S201: Identify the first voice information in the voice information of the target object, wherein the first voice information is voice information with a volume lower than a preset volume;

[0058] Step S202: Remove the first voice information from the voice information of the target object to obtain the second voice information corresponding to the target object;

[0059] Step S203: Convert the second voice information into the first text information.

[0060] Optionally, the target customer may make some low-volume sounds while conducting business, such as whispering or slight coughing. These sounds are insignificant to bank customers. Therefore, the voice information processing system will identify these voice messages with volumes lower than a preset volume as the first voice message and filter them out from the target customer's voice message, using the remaining voice message as the target customer's second voice message. Finally, the voice information processing system converts the second voice message into first text message.

[0061] In an optional embodiment, the voice information processing system can further identify whether a first type of keyword exists in the first text information, wherein the first type of keyword is used to characterize information that is prohibited from being displayed to a first object, and the first object is the object viewing the target projection device. If the first type of keyword exists in the first text information, the voice information processing system determines that the information corresponding to the first type of keyword in the first text information is information to be encrypted; if the first type of keyword does not exist in the first text information, the voice information processing system determines that there is no information to be encrypted in the first text information.

[0062] Optionally, the first category of keywords can be keywords related to the bank's internal privacy data, such as the bank's reserve amount, the location of the bank's reserve funds, the location of the bank's fund transportation, and the identity information of bank staff. The first category of keywords can also be keywords related to the privacy data of bank customers, such as the bank customer's identity information, the bank customer's deposit amount, and the bank customer's residential location.

[0063] To prevent the aforementioned internal bank privacy data and bank customer privacy data from being publicly displayed on the target screen projection device, the voice information processing system identifies the internal bank privacy data and / or bank customer privacy data that may exist in the first text information by recognizing the first type of keywords, and then encrypts this privacy data as information to be encrypted during the subsequent display process.

[0064] In one optional embodiment, after converting the target object's voice information into first text information, the voice information processing system can further detect whether a second type of keyword exists in the first text information. This second type of keyword is related to a business form, which is a form that the first object needs to fill out or approve. If the second type of keyword exists in the first text information, the voice information processing system sends the business form to the target projection device for display; if the second type of keyword does not exist in the first text information, the voice information processing system prohibits sending the business form to the target projection device for display.

[0065] Optionally, the second keyword can be related to business forms that the first user needs to fill out or review, such as "fill out," "form," "credit card application form," "bank card cancellation application form," etc. To facilitate the display of these business forms to customers, when the voice information processing system recognizes the presence of a second type of keyword in the first text information, it can automatically retrieve the form corresponding to the second type of keyword from the database and display it on the target projection device. For example, if the voice information processing system recognizes the keyword "credit card application form" in the first text information, it will automatically retrieve an electronic credit card application form and display it on the target projection device.

[0066] In one optional embodiment, after converting the target object's voice information into first text information, the voice information processing system detects whether a third type of keyword exists in the first text information. This third type of keyword is a keyword related to the target financial product, which is a financial product recommended by the target object to the first object. If the third type of keyword exists in the first text information, the voice information processing system sends the product information of the target financial product to the target projection device for display; if the third type of keyword does not exist in the first text information, the voice information processing system prohibits sending the product information of the target financial product to the target projection device for display.

[0067] Optionally, during the processing of certain business scenarios, the target audience may recommend financial products to bank customers. To more conveniently display the product information of these financial products, the voice information processing system in this application can identify third-category keywords related to the target financial product, such as the name of the target financial product. If the voice information processing system recognizes that the first text information contains keywords such as the name of the target financial product, the voice information processing system automatically reads the product information of the target financial product from the database and displays it on the target projection device.

[0068] In one optional embodiment, after sending the second text information to the target projection device for display, the voice information processing system can further acquire the voice information of the first object and identify whether a fourth type of keyword exists in the first object's voice information. This fourth type of keyword represents a question posed by the first object to the target object in response to the second text information. If the fourth type of keyword exists in the first object's voice information, the voice information processing system performs a distinguishing text operation on the target text information in the second text information. The target text information is the text information related to the fourth type of keyword, and the distinguishing text operation switches the text format of the target text information to a different format than the text formats of other text information. The other text information refers to the text information in the second text information excluding the target text information. If the fourth type of keyword does not exist in the first object's voice information, the voice information processing system prohibits performing a distinguishing text operation on the target text information in the second text information.

[0069] Optionally, when the first object is viewing the second text information on the target projection device, it may have questions about certain content. Therefore, the first object will ask the target object for these questions. In order to better facilitate the target object's understanding of the first object's questions and improve the target object's comprehension efficiency, the voice information processing system can acquire the first object's voice information and then identify whether there are fourth-type keywords in the first object's voice information. The fourth-type keywords are used to represent the questions raised by the first object to the target object regarding the second text information. For example, the fourth-type keywords can be "what does paragraph xx mean", "the meaning of paragraph xx", etc.

[0070] When the speech information processing system recognizes that the fourth type of keyword exists in the voice information of the first object, the speech information processing system will perform differential text operations on the target text information in the second text information, such as bolding, enlarging the font, and highlighting the color of the target text information.

[0071] By performing differential text operations on the target text information, the target audience can quickly understand the content that the first audience has questions by viewing the target projection device. The first audience can also re-understand the content by combining it with the target text information, thereby achieving the technical effect of improving business processing efficiency.

[0072] In one optional embodiment, the voice information system may include a first device and a second device, wherein the first device is a device used by bank customers and the second device is a device used by bank staff. Both the first device and the second device can be used as audio acquisition devices and screen projection devices.

[0073] Optionally, as shown in Figure 3, the first device includes at least Bluetooth, a display screen, and a microphone. Bluetooth ensures a one-to-one connection between the first and second devices within the same window; the display screen displays second text information converted from the target object's voice information; and the microphone collects the bank customer's voice information.

[0074] Optionally, such as Figure 4 As shown, the second device includes at least an external microphone, an external device cable, Bluetooth connectivity, and a display screen. The external microphone is used to collect voice information from bank staff (e.g., the target audience), then the processor runs a program to convert the target audience's voice information into second text information and displays it on the display screen. The external device cable is used to connect external devices such as a mouse and keyboard, allowing corrections to be made if errors occur during the voice-to-text conversion. Bluetooth connectivity ensures a one-to-one connection between the second device and the first device.

[0075] In one alternative embodiment, Figure 5A flowchart of another voice information processing method according to an embodiment of this application is shown, such as... Figure 5 As shown, it includes the following steps:

[0076] S501: Detects whether the teller device (second device) and the customer device (first device) are connected via Bluetooth to ensure that the text converted from voice can be displayed on both the teller device and the customer device at the same time, thus ensuring smooth information communication.

[0077] S502: When the external receiver of the teller device is connected to the display screen of the teller device, the display screen will show "Receiver connected to device" to ensure that the teller device can receive the teller's voice information.

[0078] S503: The teller (target) begins to speak. The voice information processing method in this application is executed through the built-in processor of the teller device to convert the teller's voice into second text information and save it in the storage module.

[0079] S504: Synchronously display the text in the storage module on the screens of the teller device and the customer device;

[0080] S505: Tellers can check whether the second text message is consistent with their intended meaning. If there is an error, they can make corrections via touchscreen or external device to avoid customer misunderstanding.

[0081] S506: Customers primarily rely on listening to the teller's voice messages, supplemented by viewing the secondary text information on the display screen. This combination of visual and auditory methods helps avoid misoperations caused by customers not being able to hear the teller's words clearly, thereby improving communication efficiency.

[0082] S507: After serving a customer, the teller device will promptly archive and save the customer's voice, the teller's voice, and the converted text to avoid future transaction disputes.

[0083] S508: If there is a financial product recommendation process, the product information of the financial product will be displayed on the display screens of the teller equipment and the customer equipment.

[0084] S509: The device refreshes the page, waiting for the next customer.

[0085] As described above, this application provides a text-based information interaction method in addition to voice interaction by converting the voice information of the target object into text information and displaying it on the target projection device corresponding to the target object. This avoids the low information interaction efficiency problem of prior art that relies solely on audio devices for information interaction. Furthermore, this application identifies the voice information of the target object from the voice information of N objects, thus avoiding displaying the voice information of other bank staff who are not associated with the target projection device on the target projection device. In addition, this application can identify whether there is information to be encrypted in the first text information. If such information exists, the first text information is encrypted, thus avoiding the poor information security problem caused by indiscriminately converting the target object's voice information into text information.

[0086] Example 2

[0087] This embodiment provides an optional voice information processing device, wherein each implementation unit / module in the voice information processing device corresponds to each implementation step in Embodiment 1.

[0088] Figure 6 This is a schematic diagram of an optional voice information processing device provided according to an embodiment of this application, such as... Figure 6 As shown, it includes: acquisition module 601, conversion module 602, recognition module 603, data processing module 604, and information sending module 605.

[0089] Specifically, the acquisition module 601 is used to acquire voice information of N objects through an audio acquisition device, and identify the voice information of the target object from the voice information of the N objects, where N is a positive integer and the target object is the object corresponding to the target projection device; the conversion module 602 is used to convert the voice information of the target object into first text information; the recognition module 603 is used to identify whether there is information to be encrypted in the first text information, where the information to be encrypted is information that is prohibited from being publicly displayed through the target projection device; the data processing module 604 is used to perform data processing operations on the first text information to obtain second text information when there is information to be encrypted in the first text information, where the data processing operation is used to encrypt the information to be encrypted in the first text information; and the information sending module 605 is used to send the second text information to the target projection device for display.

[0090] Optionally, the recognition module includes a voice feature recognition unit and a voice information determination unit. The voice feature recognition unit is used to perform voice feature recognition processing on the voice information of N objects using a target model to obtain the voice features of each object. The target model is a neural network model trained using the voice information of M objects with known voice features as training samples, where the M objects include at least N objects. The voice information determination unit is used to determine the voice information corresponding to the voice features of the target object as the voice information of the target object.

[0091] Optionally, the conversion module includes: a first recognition unit, a first voice information processing unit, and an information conversion unit. The first recognition unit is used to recognize first voice information in the voice information of the target object, wherein the first voice information is voice information with a volume lower than a preset volume; the first voice information processing unit is used to remove the first voice information from the voice information of the target object to obtain second voice information corresponding to the target object; and the information conversion unit is used to convert the second voice information into first text information.

[0092] Optionally, the identification module includes: a second identification unit, a first determination unit, and a second determination unit. The second identification unit is used to identify whether a first type of keyword exists in the first text information, wherein the first type of keyword is used to characterize information that is prohibited from being displayed to a first object, the first object being the object viewing the target projection device; the first determination unit is used to determine, if the first type of keyword exists in the first text information, that the information corresponding to the first type of keyword is the information to be encrypted; the second determination unit is used to determine, if the first type of keyword does not exist in the first text information, that there is no information to be encrypted in the first text information.

[0093] Optionally, the voice information processing device further includes: a first detection module, a first processing module, and a second processing module. The first detection module is used to detect whether a second type of keyword exists in the first text information, wherein the second type of keyword is a keyword related to a business form, and the business form is a form that needs to be filled out or approved by the first object; the first processing module is used to send the business form to the target projection device for display if the second type of keyword exists in the first text information; the second processing module is used to prevent the business form from being sent to the target projection device for display if the second type of keyword does not exist in the first text information.

[0094] Optionally, the voice information processing device further includes: a second detection module, a third processing module, and a fourth processing module. The second detection module is used to detect whether a third type of keyword exists in the first text information, wherein the third type of keyword is a keyword related to the target financial product, and the target financial product is a financial product recommended to the first target by the target object; the third processing module is used to send the product information of the target financial product to the target projection device for display if the third type of keyword exists in the first text information; the fourth processing module is used to prohibit sending the product information of the target financial product to the target projection device for display if the third type of keyword does not exist in the first text information.

[0095] Optionally, the voice information processing device further includes: an acquisition module, a first recognition module, a fifth processing module, and a sixth processing module. The acquisition module is used to acquire the voice information of a first object; the first recognition module is used to identify whether a fourth type of keyword exists in the voice information of the first object, wherein the fourth type of keyword represents a consultation question raised by the first object to the target object in response to the second text information; the fifth processing module is used to perform a distinguishing text operation on the target text information in the second text information when the fourth type of keyword exists in the voice information of the first object, wherein the target text information is text information related to the fourth type of keyword, and the distinguishing text operation is used to switch the text format of the target text information to a text format different from the text format of other text information, wherein other text information refers to text information in the second text information other than the target text information; the sixth processing module is used to prohibit the distinguishing text operation on the target text information in the second text information when the fourth type of keyword does not exist in the voice information of the first object.

[0096] Example 3

[0097] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the voice information processing method of any one of the above embodiments 1 by executing the executable instructions.

[0098] Figure 7 This is a schematic diagram of an electronic device according to an embodiment of this application, such as... Figure 7 As shown, this application provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the voice information processing method in Embodiment 1 above.

[0099] Example 4

[0100] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which includes a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the voice information processing method in Embodiment 1 above when it is running.

[0101] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0102] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0103] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0107] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for processing voice information, characterized in that, include: The audio acquisition device collects the voice information of N objects, and identifies the voice information of the target object from the voice information of the N objects, where N is a positive integer and the target object is the object corresponding to the target projection device; Convert the voice information of the target object into first text information; Identify whether there is information to be encrypted in the first text information, wherein the information to be encrypted is information that is prohibited from being publicly displayed through the target projection device; If the information to be encrypted is present in the first text information, a data processing operation is performed on the first text information to obtain the second text information, wherein the data processing operation is used to encrypt the information to be encrypted in the first text information. The second text information is sent to the target projection device for display.

2. The method according to claim 1, characterized in that, Identifying the voice information of the target object from the voice information of the N objects includes: The speech information of the N objects is processed by the target model to identify the speech features of each object. The target model is a neural network model trained using the speech information of M objects with known speech features as training samples. The M objects include at least the N objects. The speech information corresponding to the voice features of the target object is used as the speech information of the target object.

3. The method according to claim 1, characterized in that, Converting the speech information of the target object into first text information includes: Identify the first voice information in the voice information of the target object, wherein the first voice information is voice information with a volume lower than a preset volume; Remove the first voice information from the voice information of the target object to obtain the second voice information corresponding to the target object; The second voice information is converted into the first text information.

4. The method according to claim 1, characterized in that, Identifying whether there is information to be encrypted in the first text information includes: Identify whether there is a first type of keyword in the first text information, wherein the first type of keyword is used to characterize information that is prohibited from being displayed to a first object, and the first object is the object that is watching the target projection device; If the first type of keyword exists in the first text information, the information in the first text information corresponding to the first type of keyword is determined to be the information to be encrypted. If the first type of keyword does not exist in the first text information, it is determined that the information to be encrypted does not exist in the first text information.

5. The method according to claim 4, characterized in that, After converting the speech information of the target object into first text information, the method further includes: Detect whether there are second-type keywords in the first text information, wherein the second-type keywords are keywords related to business forms, and the business forms are forms that the first object needs to fill out or approve; If the second type of keyword exists in the first text information, the business form will be sent to the target projection device for display. If the second type of keyword does not exist in the first text information, the business form shall not be sent to the target projection device for display.

6. The method according to claim 4, characterized in that, After converting the speech information of the target object into first text information, the method further includes: Detect whether there are third-category keywords in the first text information, wherein the third-category keywords are keywords related to the target financial product, and the target financial product is the financial product recommended by the target object to the first object; If the third type of keyword exists in the first text information, the product information of the target financial product will be sent to the target projection device for display. If the third type of keyword does not exist in the first text information, it is prohibited to send the product information of the target financial product to the target projection device for display.

7. The method according to claim 4, characterized in that, After sending the second text information to the target projection device for display, the method further includes: Obtain the sound information of the first object; Identify whether a fourth type of keyword exists in the voice information of the first object, wherein the fourth type of keyword is used to characterize the consultation question raised by the first object to the target object in response to the second text information; If the fourth type of keyword exists in the sound information of the first object, a distinguishing text operation is performed on the target text information in the second text information. The target text information is text information related to the fourth type of keyword. The distinguishing text operation is used to switch the text format of the target text information to a text format different from the text format of other text information. The other text information is text information in the second text information other than the target text information. If the fourth type of keyword is not present in the sound information of the first object, the distinguishing text operation is prohibited on the target text information in the second text information.

8. A voice information processing device, characterized in that, include: The acquisition module is used to acquire voice information of N objects through an audio acquisition device, and to identify the voice information of a target object from the voice information of the N objects, where N is a positive integer and the target object is the object corresponding to the target projection device; The conversion module is used to convert the voice information of the target object into first text information; The identification module is used to identify whether there is information to be encrypted in the first text information, wherein the information to be encrypted is information that is prohibited from being publicly displayed through the target projection device; A data processing module is used to perform data processing operations on the first text information to obtain second text information when the information to be encrypted is present in the first text information, wherein the data processing operation is used to encrypt the information to be encrypted in the first text information. The information sending module is used to send the second text information to the target projection device for display.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the voice information processing method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the voice information processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice recognition method and device, electronic equipment, medium and program product

    CN113051895A

  • Audio processing method and device based on automatic speech recognition

    CN115482823A