System for operational analysis of communication based on the correlation between data and voice commands
The system addresses the challenge of efficiently analyzing large volumes of audio files by correlating data and voice commands using AI and data science, enabling structured search, accurate event correlation, and precise operator evaluations, thereby reducing costs and improving operational reliability.
Patent Information
- Application Number
- PCT/BR2024/050593
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
The large volume of stored audio files in supervisory systems, such as those in the electrical sector, makes it difficult to obtain information efficiently, leading to high financial costs and inaccurate operator evaluations due to the reliance on manual listening and note-taking.
A system for operational analysis of communication that correlates data and voice commands using artificial intelligence and data science techniques, enabling structured search and automatic correlation of audio files with supervisory systems, and providing comprehensive operator evaluations.
The system reduces the time required to search for information in audio files, performs accurate automatic correlations between events and recordings, and provides precise operator evaluations, thereby reducing costs and improving operational reliability.
Smart Images

Figure BR2024050593_26062025_PF_FP_ABST
Abstract
Description
"SYSTEM FOR OPERATIONAL ANALYSIS OF COMMUNICATION BASED ON CORRELATION BETWEEN DATA AND VOICE COMMANDS" FIELD OF INVENTION
[0001] The present invention relates to a system for operational analysis of communication based on the intelligent correlation of data and voice commands. More specifically, the invention relates to a speech-to-text and speech recognition system capable of identifying keywords related to a given context and to one or more supervisory systems in audio files. FUNDAMENTALS OF THE INVENTION
[0002] Currently, part of the Brazilian electrical system is remotely controlled (from generation and transmission to distribution of electrical energy), meaning that it is possible to remotely perform interventions and commands on the network from an operations center. The operation is made possible through conversations via telephone (stereo audio, with separate speakers per channel) and radio (mono audio, with both speakers on the same channel) with field teams and other agents (operating companies in the electrical sector), who verify the interventions performed remotely, as well as perform interventions and services that can only be performed in the field.These conversations are then stored for auditing purposes, to analyze compliance between what was recorded by the National System Operator (ONS) and the energy agent, to obtain more information about events that occurred in the network, to evaluate events that occurred in a work shift, to evaluate communication between operators, among others. applications .
[0003] However, the large volume of stored audio files makes it very difficult to obtain information, since the audio files need to be listened to and the information noted down manually, which results in a high financial cost, resulting from the need to allocate several people to analyze the audio files from a specific time interval of interest. Furthermore, the large amount of audio files needed to obtain information about a specific event often makes this search unfeasible, leading to a loss of opportunity on the part of the energy agent, who ends up knowing less about the events that occurred in his network.
[0004] Another problem related to the large amount of audios is the need to evaluate operators with only a few audios, which makes the evaluation dependent on a set of chosen audios, which can generate inaccurate results about the operators' conduct during the operation, since it does not take into account all the audios in which the agent acted.
[0005] It is important to note that similar problems occur in different supervisory systems that rely on telephone or radio conversations, such as mining systems or railway, waterway or air transport.
[0006] In order to solve such problems, the present invention provides an operational analysis system for communication based on the correlation between data and voice commands, aiming to enable the structured search for information in calls / audios generated from the operation and the correlation of these audios with supervisory systems. as, for example, in the electricity sector, through the creation of automatic correlations between connections and events occurring in the network, using artificial intelligence and data science techniques.
[0007] Thus, the present invention aims to reduce the time required to search for information in audio files, perform automatic correlations between events occurring in supervisory systems and recordings / calls related to them, and evaluate operators based on all audios in which they have participated as interlocutors.
[0008] Thus, the present invention aims to provide a reduction in audio audit costs, enabling the search for important information quickly and easily, and generating different views of the business that would not be possible without the invention. The invention also aims to contribute to increasing the operational reliability of the centers as it aims to increase operational safety. Furthermore, by contributing to a more precise evaluation of operators, in accordance with the regulations of the area, the invention brings significant gains to the entire sector that uses operations oriented to audio communications.
[0010] There are several studies known, in the current state of the art, that carried out basic and applied research on the subject of audio-to-text transcription. However, there are no solutions that consist of a functional real-time system applied capable of solving the mentioned problems related to supervisory systems.
[0011] Conversion applications are also known. speech-to-text and speech recognition in some distinct cases such as, for example: operation and control centers, call centers, conferences and other types of devices with this type of conversion. 0 which indicates a tendency towards the use of this type of technique in cases where there is voice communication of relevant operational information.
[0012] Some documents from the current state of the art will be presented as a way of illustrating the current state of technology.
[0013] The paper "CallSurf - Automated transcription, indexing and structuring of call center conversational speech for knowledge extraction and query by content" presents an analysis of text transcription and data mining systems that describes that call centers contain a huge amount of information of all types in the form of conversational speech. According to this paper, an efficient way to explore unstructured data from phone calls is data mining, but the techniques apply only to text. So, the paper presents a complete platform, from automatic transcription to information retrieval and data mining.
[0014] Document US7788095B2 describes a method and apparatus for indexing one or more audio signals using a speech-to-text engine and a phoneme detection engine and generating a combined network comprising a text part and a phoneme part. A word to be searched for is searched for in the text part and, if not found or found with low index, certainly, it is divided into phonemes and searched in the phoneme parts of the network. However, the method described is applied specifically to call-center monitoring.
[0015] Document US20210304107A1 discloses a system including an audio input device, a transmitter device, a gateway device, and a server computer. The audio input device may be configured to capture audio. The transmitter device may be configured to receive audio from the audio input device and communicate the audio wirelessly. The gateway device may be configured to receive audio from the transmitter device and generate an audio stream in response to preprocessing the audio. The server computer may be configured to receive the audio stream, execute computer-readable instructions that implement an audio processing engine, and provide a report in response to the audio stream.The audio processing engine may be configured to distinguish between a plurality of voices in the audio stream, convert the plurality of voices into a text transcript, perform analytics on the audio stream to determine metrics, and generate the report based on the metrics.
[0016] As can be seen from the examples of prior art documents, several speech-to-text conversion mechanisms with specific term identification steps are currently known. However, no known technology offers a structured search for information in calls / audios generated from the operation and the performance of automatic correlations between the calls and the events that occurred during them. the operation determined from the application of artificial intelligence and data science techniques. Thus, the state of the art does not present a solution capable of facilitating the search for post-operation calls while correlating them with events recorded in operating and supervisory systems, providing an operator with a complete view of the events of each call.
[0017] As will be further detailed below, the present invention aims to solve the problems of the prior art described above in a practical and efficient manner. SUMMARY OF THE INVENTION
[0018] The present invention aims to provide operational analysis of communication based on the intelligent correlation of data and voice commands, aiming to enable the structured search for information in calls / audio files generated from the operation and the realization of automatic correlations between calls and events that occur based on the application of artificial intelligence and data science techniques.
[0019] In order to achieve the objectives described above, the present invention provides a system for operational analysis of communication from the intelligent correlation of data and voice commands comprising: a cloud audio transcription module configured to: load audio files and audio metadata of telephone or radio calls of the operation from a database; perform digital pre-processing of signals in the audio files; transcribe the audio files; and store each transcription and audio metadata in a database with a search and analysis mechanism distributed and optimized data for texts; and additionally comprising: a correlation module configured to correlate events that have occurred, present in the agent's supervisory and ERP systems, as well as other relevant external systems, with the transcripts. In these correlations, the following are used as input: a list of events from the supervisory, EAM and external systems that occurred in a given time interval, the transcripts and the audio metadata; and said correlation module additionally configured to correlate transcripts with each other, in which the following are used as input: the transcripts, the audio metadata and transcripts correlated with events that have occurred.
[0020] The system according to the present invention consists in the fact that the audio metadata comprises the date, time and duration of the call, and the origin and destination extension.
[0021] Furthermore, the system according to the present invention consists in the fact that the correlated transcripts are compiled in a web application and in communication search and analysis modules.
[0022] Furthermore, the system according to the present invention consists of the fact that the pre-processing comprises: channel separation if necessary, removal of silence, digital signal processing techniques and when applicable, application of an algorithm for joining the audios with channel separation and structuring of the transcription separating the interlocutors.
[0023] Additionally, the system according to the present invention consists in the fact that the correlation between events that occurred, present in the supervisory systems, EAM and external systems, with the transcriptions of the calls comprises: extracting information from the transcriptions by means of natural language processing techniques; comparing transcription to transcription with the list of events that occurred, from the time the call was made, based on the audio metadata and based on the list of events that occurred; and if the recognized entities of the transcription are identified in the list of events that occurred, then the event that occurred is correlated to the transcription.
[0024] The system according to the present invention also consists in the fact that the correlation of transcripts among themselves comprises: comparing transcripts, within a defined time interval, that deal with the same equipment, documents or service orders; and grouping these transcripts, so that they form part of a history.
[0025] Furthermore, the system according to the present invention consists in the fact that a training for the classification of calls comprises: manually classifying the transcripts with classes of events that occurred defined by the agent; after the manual classification of the transcripts, training an intelligent classifier model in a chain to learn to recognize the events that occurred in each manually classified transcript; after training the chain intelligent classifier model, said classifier model retroactively evaluates possible transcripts that were not manually classified, conferring one or more classes to said transcripts that were not manually classified.
[0026] Furthermore, the system according to the present invention consists in the fact that the intelligent chain classifier model allows a transcript to assume two or more classes and establishes relationships between events that have occurred.
[0027] Additionally, the system according to the present invention consists in the fact that the retraining of the intelligent chain classifier model is performed as new transcripts are manually classified and as possible classification errors of the classifier model are manually corrected; in which when new transcripts are manually classified, the model restarts the training process using the new data, in which: if the metrics of the retrained model are better than those of the original model, the retrained model begins to be used as the official model; and otherwise, the original model is used as the official model.
[0028] The system according to the present invention comprises an operator evaluation module configured to apply natural language processing algorithms to the transcripts in conjunction with machine learning to assess a score, based on criteria defined by the agent, for each telephone call of each operator, in that the total score is the average of the scores of all the operator's calls given by the average of the criteria.
[0029] Additionally, the system according to the present invention has a shift exchange module, where the most relevant information occurring during the shift is identified and structured in an interface dedicated to the synthesis of this content for the understanding of other users.
[0030] Additionally, the system according to the present invention has an occurrence analysis module, dedicated to the energy transmission segment, where specific models are executed to synthesize information on relevant events that were correlated with other systems (Supervisory, ERP and third-party systems) in a customized interface to facilitate the analysis and investigation of occurrences related to occurrences in transmission lines, mainly when there is unavailability of the same.
[0031] Additionally, the system according to the present invention has a generation limitation module, dedicated to the energy generation segment, where specific models are executed to synthesize information on relevant events that were correlated with other systems (Supervisory, ERP and third-party systems) in a customized interface to facilitate the analysis and investigation of occurrences that concern energy generation limitations in power plants.
[0032] Additionally, the system according to the present invention has a state change module, dedicated to the energy generation segment, where specific models were executed to synthesize information on relevant events that were correlated with other systems (Supervisory, ERP and third-party systems) in a customized interface to facilitate the analysis and investigation of occurrences related to changes in the status of generating units of power plants.
[0033] Furthermore, the system according to the present invention consists in the fact that the transcription of the audio files by an audio-to-text transcription model uses customized vocabulary that contains the context of the operating sector.
[0034] Additionally, the system according to the present invention consists of the fact that in the transcription of audios by an audio-to-text transcription model, the transcription model used presents a word error rate of less than 0.5 and allows the insertion of personalized vocabulary. BRIEF DESCRIPTION OF THE FIGURES
[0035] The detailed description presented below makes reference to the attached figures and their respective reference numbers.
[0036] Figure 1 illustrates a general schematic diagram of the system of the present invention.
[0037] Figure 2 illustrates the audio pre-processing methodology that precedes transcription.
[0038] Figure 3 illustrates a transcript correlation model with other systems according to the present invention.
[0039] Figure 4 illustrates a model for correlating transcripts with each other according to the present invention.
[0040] Figure 5 illustrates a training for classifying connections according to the present invention.
[0041] Figure 6 illustrates the classification model, already trained, acting in accordance with the present invention.
[0042] Figure 7 illustrates the operator evaluation module as proposed by the present invention.
[0043] Figure 8 illustrates the interface modules existing in the system. DETAILED DESCRIPTION OF THE INVENTION
[0044] As a preliminary point, it should be noted that the following description will be based on a preferred embodiment of the invention based on an operation and post-operation support tool that allows the search for recorded calls (the term "calls" refers to both telephone calls and radio calls in operation) and provides the correlation of these calls with supervisory systems and the correlation between calls (call histories) by means of the correlation of transcripts with operating systems and the correlation between transcripts, respectively. To this end, the system supports qualitative analyses in the communication carried out by the operators of the operation centers and allows the traceability of the events that occurred and mentioned in the calls. In addition, it is important to note that the description that follows refers to the use of the present invention in an operation center in the electric sector.As will be evident to anyone skilled in the art, however, the invention is not limited to this particular embodiment.
[0045] Figure 1 illustrates a general schematic diagram of the system according to the present invention, comprising a cloud audio transcription module 10, said module transcription module 10 adapted to load audio files 11 of telephone or radio calls from the operation, and audio metadata 12 (date, time and duration of the call, origin and destination extension, among other information related to the call itself), of telephone or radio calls, initially stored in a database. In the transcription module 10, the audio files 11 are subjected to pre-processing 13 that includes: channel separation if necessary, removal of silence, digital signal processing techniques and, when applicable, application of an algorithm to join the audios with channel separation and structuring of the transcription separating the interlocutors. The audio files 11 contain crucial information about events that occurred during the operation, including requests for actions in the field, execution of service orders, communication with external agents, provision of assets, adjustments of asset parameters and safety documents.These varied actions are monitored in multiple operating systems 20, such as SCADA, EAM and ONS systems (SATRA, SAGER, SGI).
[0046] Still in the transcription module 10, after loading the audio files 11 and pre-processing 13, the audio files 11 are transcribed 14 by converting the audio files 11 into text format files (also called transcripts). Each transcript 14 is subsequently stored in a database 30. As will be apparent to one skilled in the art, the transcription 14, performed with cloud provider APIs, can be implemented in different ways within the context of the present invention, both with APIs from the main cloud providers (AWS, Google Cloud Platform or Azure) and with locally developed and trained transcribers.
[0047] The transcripts 14, together with the audio metadata 12, said metadata 12 comprising information related to the audio files 11, are then stored in the database 30 with a distributed and text-optimized data search and analysis engine. There are different search and data analysis engines that can be used, such as Elastí cSearch, Solr, ArangoDB, Vespa, Algolía, among others, so this does not represent a limitation to the scope of the invention, and other models can therefore be adopted.
[0048] Additionally, the system comprises an artificial intelligence module 40 with specialist machine learning and natural language processing algorithms, said artificial intelligence module 40 comprising correlation algorithms 41 and automatic link classification 42.
[0049] The portion of the module that includes the correlation algorithms 41 is adapted to correlate the information on events that have occurred, present in the agent's operating systems 20, with the transcripts 14 stored in the database 30 (illustrated in detail in Figure 3), in which the following are used as input: a list of detailed events from the operating systems that occurred in a given time interval 200, the transcripts 14 and the audio metadata 12. Thus, the correlation algorithms 41 are responsible for automatically correlating the information on events that have occurred, coming from various operating systems 20 of the agent, with the transcripts 14. Correlation is one of the main characteristics of this invention, since from information on events that occurred, extracted from the transcripts 14 and the audio metadata 12, it is possible to associate the transcripts 14 with events that occurred in these operating systems 20. Thus, the present invention considers unique information from each call, such as call number, person's name, equipment identifier, changes in equipment status, locations, facilities or identification document. The interaction of each call is important, so that equipment identifiers, numbers and identifications of events that occurred mentioned in the call are important for the proposed solution. From these correlations, post-operation can generate reports and perform analyses more easily.To this end, preferably, the correlation module 40 uses as input: the list of events that occurred in the supervisory system 200 and the transcripts 14 stored in the database 30 with a distributed and optimized data search and analysis mechanism for texts.
[0050] In addition to the correlation algorithms 41, the artificial intelligence module 40 also has the classification of links 42 (illustrated in detail in Figures 5 and 6), in which the following are used as input: the transcripts 14 and the audio metadata 12. Using natural language processing and artificial intelligence techniques, the transcripts are separated by subject through supervised learning methods.
[0051] Therefore, the transcripts 14 can be subjected to transcript classification algorithms 42, correlation 41 of transcripts 14 among themselves and correlation of transcripts 14 with operating systems 20. Optionally, the same transcripts 14 are also used to evaluate operators 43 and to summarize events that occurred in a given time interval. The correlated transcripts 14 and the results are compiled in a web application 50 that allows the search and analysis of calls, as well as in communication search and analysis modules 60, such as PowerBI, that use this data to generate visualizations of indicators for management.
[0052] Figure 2 illustrates in detail the steps performed by the cloud audio transcription module 10 as proposed by the present invention, in which it is foreseen that, initially, the audio files 11, originating from recordings made over calls between telephone extensions, stereo audios Ila (dual-channel), or radio, mono audios 11b (single-channel) undergo pre-processing 13. These audios 11a, 11b undergo different processes before being transcribed to ensure the separation of the interlocutors.
[0053] The stereo audios 11a are subjected to channel separation 131, which separates the speakers 132 before removing silence 133 in the two audios resulting from the previous process. In the same way, the mono audio 11b goes through the silence removal step 133 directly. For mono audio files 11b, this can occur through the application of a separation model, such as SepFormer, or through a transcription 14 with diarization, in which the transcription 14 itself identifies the speakers. If the audio file is stereo 11a, the speakers are separated through a channel division, that is, one speaker is considered per audio channel and vector separation is performed. To this end, as is known, the premise is that in stereo audio files 11a there is only one speaker per channel. After speaker separation 132, preferably, the audios undergo a silence removal process 133, to reduce the computational cost required to perform transcription, since the cost is proportional to the duration of the audio.
[0054] The audios with silence removed are then subjected to digital signal processing techniques 134 with the aim of improving the quality of the audio that will be transcribed 14.
[0055] Subsequent to the digital pre-processing of signals 134 of the audio files 11a, 11b, said audio files 11a, 11b are then transcribed 14 by an audio-to-text transcription model, also called speech-to-text. The transcription 14, in turn, is also customized and receives retraining based on some factors 141, namely: context of the operation, keywords, common phrases in the operation, audios of the operation and manual transcriptions. In this way, the transcription 14 ends up performing better for the recordings of the sector and uses the context of the field of application of the supervisory system 20. In this sense, it is emphasized that the present invention can be applied in different technical fields to evaluate the communication carried out in sectors such as the electric power sector, mining, waterway, rail and air transport, among others.For example, in the case of the electric power sector, the customized vocabularies used must be suitable to increase the probability of. identification of words that are most common in the context of the electrical sector, such as transformer, reactor, among others. Therefore, it will be clear to a technician in the subject the need to adjust the personalized vocabulary used according to the application.
[0056] Most preferably, any audio-to-text transcription models can be used, as long as they have a word error rate (WER) of less than 0.5, allow for the insertion of personalized vocabulary, and have a keyword recognition percentage greater than 60%.
[0057] Finally, for mono audios 11b, the result of the transcription model 14 is saved in the database 30 with a distributed and optimized data search and analysis engine for texts as already mentioned above; and for stereo audios 11a, an algorithm is applied to join the previously separated audios and to synchronize the transcription 134 of each of them. With the unified transcription 14, it is then stored in the database 30.
[0058] Figure 3 illustrates an exemplary model of the algorithms for correlating links with operating systems 41a, present in the correlation algorithm 41, correlating transcripts 14 with events that occurred in operating systems 20, in which the following are used as input: a list of detailed events of each operating system that occurred in a certain time interval 200, the transcripts 14 and the audio metadata 12. Information is extracted from the transcripts 14, through natural language processing techniques (Natural Language Processing - NLP), and techniques of recognition of entities, such as proper names, locations, equipment, facilities, dates, numbers, times, among others. With the recognized entities, transcript 14 is compared with the list of events that occurred 200, from the time the call was made, based on the audio metadata 12 and based on the list of events that occurred 200. Among the information analyzed in transcript 14 are, for example, the date or time cited, the locations, numbers that may refer to specific service orders, identification of equipment, documents, etc. If the recognized entities in transcript 14 are identified in the list of events that occurred 200, the event that occurred is then correlated with transcript 201.
[0059] Figure 4 illustrates a model for correlating transcripts 41b, present in the correlation algorithm 41, in which the following are used as input: transcripts 14, audio metadata 12 and transcripts correlated to events that occurred 201. In this model 41b, transcripts 14 within a defined time interval that deal with the same equipment, documents, interventions and / or service orders are compared. These transcripts 14 are grouped and considered part of a story .
[0060] In this way, the present invention achieves the objective of facilitating the search for calls for post-operation while correlating the transcripts 14 of the calls to events that occurred in the operating system 20, providing the operator with a complete view of the event that the call was about.
[0061] Additionally, the system according to the present invention may present additional features that assist in the analysis of connections.
[0062] Figure 5 illustrates the training for the classification of calls 42. Initially, it is necessary that the transcripts are manually classified 421 into different types (or labels), in the context of power generation, one can mention "Voltage Adjustment", "Generation Modulation" and "Change of State", so that a transcript can be classified with more than one label. After the manual classification of the transcripts 421, a supervised and in-chain multi-label machine learning model 43 is trained to learn to automatically recognize the labels of events occurring in each transcript 14.
[0063] This intelligent classifier model 43 allows a transcript 14 to have more than one label, that is, a transcript 14 can deal with one or more event labels. Furthermore, because it consists of a chain model, it establishes relationships between the types of events that occurred, increasing the probability of identifying related events in the transcript.
[0064] Figure 6 illustrates the classification model 43 - already trained - acting in the automatic identification of event labels in the transcripts 14 from supervised and chained machine learning.
[0065] After training, the intelligent classifier model 43 evaluates the transcripts that were not used for training (manually labeled transcripts), assigning one or more event types to each one of them.
[0066] Because it consists of a chain model, the classifier model 43 ends up consisting of several models (one for each type of event) that have as output the occurrence - or not - of the given event in the transcript 14. In this way, the model 43 is able to use the results of the classifications of each type of event individually for the next event. Therefore, the output of the model for the first type of event 422, which has as input the transcript 14 and the link metadata 12, is one of the inputs in the model of the second type of event 433; the model of the third type of event relies on the results of the first and second models, and so on. The one or more classes are then saved in the database with a distributed and text-optimized data search and analysis engine 30.
[0067] To improve the accuracy of the intelligent classifier model 43, retraining of said classifier model 43 is performed as new transcripts are manually classified and as possible classification errors of the classifier model 43 are manually corrected. Thus, when new transcripts are manually labeled, the classifier model 43 restarts the training process using the new data, where: if the metrics of the retrained model 43 are better than those of the original model 43, the retrained model 43 begins to be used as the official model 43; and otherwise, the original model 43 is used as the official model 43.
[0068] Figure 7 illustrates the operator evaluation module according to the present invention. This optional module has the function of evaluating whether the operator's communication is adherent to the operation's internal procedures. To this end, natural language processing (NLP) algorithms are applied to the transcripts 14 in conjunction with machine learning to confirm whether or not the criterion was met (0 or 100) in pre-defined criteria by the operation for each telephone call of each operator, with the total score being the average of the scores of all the operator's calls, which in turn is given by the average of the criteria.
[0069] Among the predefined criteria for the operator evaluation module, the following stand out: identification (operator, operation center, actions, equipment, locations); use of polite words; use of vulgar words; confirmation (actions, equipment, locations, safety); assertiveness. There are also evaluation metrics such as density of technical terms, speech speed and number of calls.
[0070] From the above, in summary, the present invention provides a system - a tool for operational analysis of communication based on the intelligent correlation of data and voice commands capable of transcribing audio recordings of the operation of a system, inserting the context of the respective sector into the transcriber through personalized vocabularies. These transcriptions are then stored in a database optimized for text search, such as Elastí cSearch, so that they can be used in the modules that make up the system, namely: • correlation: capable of correlating events that occurred with transcripts using data from supervisory systems, as well as relating links between them. • call classification: capable of classifying transcripts into types of events relevant to the operation and capable of retraining the machine learning model by inserting new labels • operator evaluation: capable of evaluating operator communication through the application of natural language processing and artificial intelligence techniques.
[0071] Figure 8 details how the results of these artificial intelligence modules are made available for viewing by the system's end users in the web application 50. The interfaces that make up the system are: Search; Operator Evaluation; Shift Change; Occurrence Analysis; Generation Limitation; State Change and System Settings.
[0072] Each of these interfaces has specific business rules, performs necessary calculations and searches for specific information in databases. This is intended to allow the user to perform various analyses based on the specific visualization of each module, synthesizing the information and facilitating its understanding in each context.
[0073] Since information about communication in operation centers (audio communication, SCADA system data, records made in EAM systems, records made by the ONS in external systems) is already centralized, these interfaces are based on modern UI / UX techniques to make this content available in the most educational way possible.
[0074] By integrating multiple systems and harvesting benefits of all of them through the application of natural language and machine learning techniques, the system presents itself as innovative for use in different areas, integrating information from different supervisory systems.
[0075] Thus, the present invention provides a system capable of facilitating the search for calls for operation and post-operation while correlating them with events recorded in operating and supervisory systems, providing an operator with a complete view of the events of each call. In addition, the present invention also allows the evaluation of the operator's communication protocol and the correlation of related calls among themselves.
[0076] In addition to the embodiments presented above, the same inventive concept may be applied to other alternatives or possibilities of using the invention.
[0077] While the present invention has been described with respect to certain preferred embodiments, it should be understood that it is not intended to limit the invention to those particular embodiments. Rather, it is intended to encompass all possible alternatives, modifications, and equivalencies within the spirit and scope of the invention as defined by the appended claims.
Claims
CLAIMS 1. System for operational analysis of communication based on the correlation between data and voice commands, comprising: a cloud-based audio transcription module (10) configured to: load audio files (11) and audio metadata (12) of telephone or radio calls from the operation from a database; perform pre-processing (13) digital signals in audio files; transcribe (14) the audio files (11); and store each transcription (14) and the audio metadata (12) in a database (30) with a distributed and optimized data search and analysis mechanism for texts; and characterized by the fact that it additionally comprises: a correlation module (40) configured to correlate events that have occurred, present in the supervisory systems (20) of the agent, with the transcriptions (14), in which the following are used as input: a list of events of the supervisory system that have occurred in a certain time interval (200), the transcriptions (14) and the metadata of the audios (12); and said correlation module (40) additionally configured to correlate transcriptions with each other, in which the following are used as input: the transcriptions (14), the metadata of the audios (12) and transcriptions correlated to events that have occurred (201).
2. System, according to claim 1, characterized by the fact that the audio metadata (12) comprises date, time and duration of the call, and extension of origin and destination.
3. System, according to claim 1, characterized by the fact that the correlated transcripts (14) are compiled in a web application (50) and in communication search and analysis modules (60).
4. System, according to claim 1, characterized by the fact that the pre-processing (13) comprises: separation of channels if necessary, removal of silence, digital signal processing techniques and when applicable, application of an algorithm for joining the audios with separation of channels and structuring of the transcription separating the interlocutors.
5. System, according to claim 1, characterized by the fact that the correlation between occurred events, present in the agent's supervisory systems (20), with the transcriptions (14) of the calls comprises: extracting information from the transcriptions (14) by means of natural language processing techniques; comparing transcription (14) to the transcription (14) with the list of occurred events (200), from the time at which the call was made, based on the audio metadata (12) and based on the list of occurred events (200); and if the recognized entities of the transcription (14) are identified in the list of occurred events (200), then the occurred event is correlated to the transcription (201).
6. System, according to claim 1, characterized by the fact that the correlation of transcripts (14) among themselves comprises: comparing transcripts (14), within a defined time interval, that deal with the same equipment, documents or service orders; and group these transcripts (14) so that they form part of a story.
7. System, according to claim 1, characterized by the fact that it additionally comprises a training for the classification of links (42): manually classifying (421) the transcripts with classes of occurred events defined by the agent; after the manual classification of the transcripts (421), training an intelligent chain classifier model (43) to learn to recognize the events occurred in each manually classified transcript (421); after the training of the intelligent chain classifier model (43), said classifier model (43) retroactively evaluates possible transcripts that were not manually classified (421), conferring one or more classes to said transcripts that were not manually classified (421).
8. System, according to claim 7, characterized by the fact that the intelligent chain classifier model (43) allows a transcription (14) to assume two or more classes and establishes relationships between events that occurred.
9. System according to claim 7 or 8, characterized in that the retraining of the intelligent chain classifier model (43) is performed as new transcripts are manually classified and as possible classification errors of the classifier model (43) are manually corrected; wherein when new transcripts are classified manually, the model (43) restarts the training process using the new data, in which: if the metrics of the retrained model (43) are better than those of the original model (43), the retrained model (43) begins to be used as the official model (43); and otherwise, the original model (42) is used as the official model (43). 10 System, according to claim 1, characterized by the fact that it comprises an operator evaluation module configured to apply natural language processing algorithms to the transcripts in conjunction with machine learning to assess a score from 0 to 100, based on criteria defined by the agent, for each telephone call of each operator, in which the total score is the average of the scores of all the operator's calls given by the average of the criteria.
11. System, according to claim 1, characterized by the fact that the transcription (14) of the audio files (11) by an audio-to-text transcription model uses customized vocabulary that contains the context of the operating sector.
12. System, according to claim 1, characterized by the fact that in the transcription of audios by an audio-to-text transcription model, the transcription model used presents a word error rate of less than 0.5 and that it allows the insertion of personalized vocabulary.
Citation Information
Patent Citations
Machine learning method and device for quickly improving text classification performance
CN110263173A
Problem generation method and device
CN112163405A
Systems and methods for providing searchable customer call indexes
US10872068B2