Conference Information Processing Method, Apparatus, Electronic Device, and Storage Medium
By calculating the difference between the size of the conference recording file and the theoretical size, aligning the conference recording file and the speech recognition text, the problem of poor alignment accuracy in the prior art is solved, and higher alignment accuracy and usage effect is achieved.
Patent Information
- Application Number
- CN202111135941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-09-27
AI Technical Summary
In the prior art, the alignment of conference recording files and speech recognition text is poor, which affects the subsequent use effect.
By obtaining the audio stream of clients that are connected to the conference, mixing them into single-channel audio and writing them to the conference recording file, the speech recognition results are obtained for each client, and by calculating the difference between the size and theoretical size of the conference recording file, the text in the speech recognition results are aligned with the conference recording file.
The accuracy of the alignment results of conference recording files and speech recognition text is improved, and the lag of conference recordings is taken into account, thereby improving the accuracy of alignment and usage effect.
Smart Images

Figure CN113971955B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to methods, devices, electronic devices, and storage media for processing conference information in the fields of intelligent voice and natural language processing, etc. Background Art
[0002] For audio conferences or audio-visual conferences, conference recording files can be generated through audio recording technology, and speech recognition texts (conference recognition texts) can be obtained through speech recognition technology.
[0003] In some cases, it is also necessary to associate the conference recording file with the speech recognition text, that is, to achieve the alignment of the conference recording file and the speech recognition text, so as to dynamically display the corresponding speech recognition text in real time when playing the conference recording, or to quickly locate to the corresponding position of the conference recording through the speech recognition text, etc. However, the accuracy of the current alignment methods is usually poor, thus affecting the subsequent usage effects, etc. Summary of the Invention
[0004] The present disclosure provides methods, devices, electronic devices, and storage media for processing conference information.
[0005] A method for processing conference information includes:
[0006] Obtaining an audio stream of a client accessing a conference;
[0007] Mixing the audio streams of different clients into a single-channel audio and writing it into a conference recording file;
[0008] For any client, performing the following processing respectively: obtaining the speech recognition result of the audio stream of the client, and obtaining the difference between the size of the conference recording file and the theoretical size, and aligning the speech recognition text in the speech recognition result with the conference recording file according to the difference.
[0009] A device for processing conference information includes: an obtaining module, a mixing module, and an alignment module;
[0010] The obtaining module is configured to obtain an audio stream of a client accessing a conference;
[0011] The mixing module is configured to mix the audio streams of different clients into a single-channel audio and write it into a conference recording file;
[0012] The alignment module is configured to, for any client, perform the following processing respectively: obtaining the speech recognition result of the audio stream of the client, and obtaining the difference between the size of the conference recording file and the theoretical size, and aligning the speech recognition text in the speech recognition result with the conference recording file according to the difference.
[0013] An electronic device, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0017] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above.
[0018] A computer program product, comprising a computer program / instructions which, when executed by a processor, implement the method as described above.
[0019] One embodiment in the above disclosure has the following advantages or beneficial effects: During the progress of a meeting, a meeting recording file can be generated based on the acquired audio stream of the client and the corresponding speech recognition result can be obtained. Moreover, for each speech recognition result acquired for each client each time, the difference between the size of the current meeting recording file and the theoretical size can be obtained, and the speech recognition text in the speech recognition result can be aligned with the meeting recording file according to the difference, that is, the lag of the meeting recording is considered during alignment, thereby improving the accuracy of the alignment result, etc.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0022] Figure 1 is a flowchart of the first embodiment of the meeting information processing method described in the present disclosure;
[0023] Figure 2 is a schematic diagram of the recording and speech recognition methods in an existing meeting;
[0024] Figure 3 is a schematic diagram of the relationship between the current time, the start time of the meeting, and the meeting recording file described in the present disclosure;
[0025] Figure 4 is a schematic diagram of the speech recognition result corresponding to a certain client described in the present disclosure;
[0026] Figure 5 This is a flowchart of the second embodiment of the conference information processing method described in the present disclosure;
[0027] Figure 6 This is a schematic structural diagram of the composition of Embodiment 600 of the conference information processing apparatus described in the present disclosure;
[0028] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement the embodiments of the present disclosure is shown. Detailed implementation manners
[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] In addition, it should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0031] Figure 1 This is a flowchart of the first embodiment of the conference information processing method described in the present disclosure. As Figure 1 shown, it includes the following specific implementation manners.
[0032] In step 101, an audio stream of a client accessing the conference is obtained.
[0033] In step 102, the audio streams of different clients are mixed into a single-channel audio and written into the conference recording file.
[0034] In step 103, for any client, the following processes are respectively performed: obtaining the speech recognition result of the audio stream of the client, and obtaining the difference between the size of the conference recording file and the theoretical size, and aligning the speech recognition text in the speech recognition result with the conference recording file according to the difference.
[0035] It can be seen that in the solution described in the above method embodiment, during the progress of a meeting, a meeting recording file can be generated based on the obtained audio stream of the client and the corresponding speech recognition result can be obtained. In practical applications, due to reasons such as the network, the audio streams at the same moment on different clients cannot arrive completely simultaneously. Therefore, before mixing the audio streams of different clients into a single-channel audio, a certain caching process will be performed on each audio stream, and the mixing will only be carried out when a certain condition is met. That is to say, there will be a certain lag in the mixing. Correspondingly, in the solution described in the above method embodiment, for the speech recognition result obtained for each client each time, the difference between the size of the current meeting recording file and the theoretical size can be first obtained, and then the speech recognition text in the speech recognition result can be aligned with the meeting recording file according to the difference. That is, the lag of the meeting recording is considered during the alignment, thereby improving the accuracy of the processing result, etc.
[0036] Figure 2 It is a schematic diagram of the recording and speech recognition methods in existing meetings. As Figure 2 shown, in a meeting, there are usually multiple clients. The system (meeting system) receives the uplink audio streams of each client and forwards them to other clients in the meeting, so that each meeting participant can hear the voices of other participants. In addition, the system can also send the audio streams of each client to an automatic speech recognition system (ASR, Automatic Speech Recognition) for speech recognition, and can obtain the speech recognition result returned by the ASR, which may include the speech recognition text, etc. In addition, the system will also mix the audio streams of different clients into a single-channel audio and write it into the meeting recording file, as Figure 2 shown, where the solid line represents the audio stream and the dashed line represents the speech recognition text. The meeting recording and speech recognition are carried out simultaneously and have the same life cycle as the meeting.
[0037] In the solution described in the present disclosure, after the meeting starts, with the access of each client, the audio streams of each client can be obtained, and the audio streams of different clients can be mixed into a single-channel audio and written into the meeting recording file. Generally speaking, for the obtained audio stream, it will first be decoded into the Pulse Code Modulation (PCM) format, and then speech recognition and mixing will be performed based on the decoded data, that is, the audio streams of different clients will be mixed into a single-channel audio.
[0038] There is no limitation on how to perform the mixing. For example, various existing mixing algorithms can be used.
[0039] As described above, due to reasons such as the network, the audio streams at the same moment on different clients cannot arrive completely simultaneously. Therefore, the mixing algorithm will perform a certain caching process on each audio stream and will only perform a mixing when a certain condition is met. The specific condition depends on different mixing algorithms. Correspondingly, there will be a certain lag in the conference recording, which also brings difficulties to the alignment of the conference recording file and the speech recognition text.
[0040] To address the above problems, in the solution described in this disclosure, a variable pcm_size can be used to record the size of the current conference recording file. By using the size of the current conference recording file and the start time of the conference, etc., the difference between the size of the current conference recording file and the theoretical size can be determined.
[0041] Correspondingly, at the start of the conference, the start time of the conference can be recorded. The start time of the conference usually refers to the absolute time when the conference starts, and the unit can be milliseconds (ms). The units of subsequent times are the same and will not be elaborated further.
[0042] In an embodiment of this disclosure, the first difference between the current time and the start time of the conference can be obtained, and the first product between the first difference and a preset constant can be obtained. The constant represents the file size of audio recording per millisecond. Furthermore, the second difference between the first product and the size of the current conference recording file can be obtained, and the second difference is used as the difference between the size of the current conference recording file and the theoretical size.
[0043] That is: pcm_offset = (current_time – conf_start_time) * K – pcm_size; (1)
[0044] Wherein, current_time represents the current time, conf_start_time represents the start time of the conference, K is a constant representing the file size of audio recording per millisecond, and the specific value can depend on the audio sampling rate. For example, when the sampling rate is 16000, K can be 32. pcm_size represents the size of the current conference recording file, and pcm_offset represents the difference between the size of the current conference recording file and the theoretical size.
[0045] Based on the above introduction, Figure 3 is a schematic diagram of the relationship between the current time, the start time of the conference, and the conference recording file described in this disclosure. As Figure 3 shown, the different-length small rectangles therein respectively represent each single-channel audio, that is, the audio streams of each client.
[0046] Through the above processing, the difference between the size of the current conference recording file and the theoretical size can be accurately and efficiently calculated, laying a good foundation for subsequent processing.
[0047] Correspondingly, for each client, the following processing can also be performed separately: obtain the speech recognition result of the audio stream of the client, and obtain the difference between the size of the current conference recording file and the theoretical size, and align the speech recognition text in the speech recognition result with the conference recording file according to the difference.
[0048] In one embodiment of the present disclosure, for each client, a long connection with the ASR can be established separately for the client, and the audio stream of the client can be sent to the ASR through the long connection, and the speech recognition result returned by the ASR can be obtained.
[0049] Preferably, the long connection can be a network socket (websocket) long connection. Each client corresponds to its own long connection respectively, so that the speech recognition of each client does not interfere with each other, ensuring the accuracy of the speech recognition result, and enabling the system to distinguish the clients corresponding to different speech recognition results, etc.
[0050] For each client, the establishment time of the corresponding long connection can also be recorded separately. During the process of continuously pushing the audio stream to the ASR, the ASR will continuously return the speech recognition result through the websocket callback.
[0051] In one embodiment of the present disclosure, in addition to the speech recognition text, the speech recognition result may further include the time stamp of the alignment object in the speech recognition text. The time stamp is the time offset of the start of the audio corresponding to the alignment object relative to the establishment time of the long connection. The alignment object includes at least one of the following: sentence, word, and phrase.
[0052] That is to say, the alignment object can include one, any two, or all three of sentences, words, and phrases. That is, alignment can be performed at the granularity of sentences, at the granularity of words, or at the granularity of phrases. It is also possible to achieve alignment between sentences and words, between sentences and phrases, or between words and phrases. In addition, it is also possible to achieve alignment between sentences, words, and phrases at the same time, which is very flexible and convenient.
[0053] The following takes the time stamps of sentences and words as an example for illustration. For phrases, the time stamp of the phrase can be determined according to the time stamps of the words. For example, if a phrase consists of two words, then the time stamp of the first word is the time stamp of the phrase.
[0054] Figure 4 It is a schematic diagram of the speech recognition result corresponding to a certain client described in the present disclosure. As Figure 4As shown, the content in the rectangular box is the speech content in the audio stream of the client, including "performing audio-to-text conversion while recording", "achieving synchronization of speech and text", and "automatically recording meeting information", etc. That is, it is assumed that the user uttered the above speech content in sequence.
[0055] As Figure 4 shown, for each speech content, the corresponding speech recognition result can be obtained respectively. For example, for "performing audio-to-text conversion while recording", the corresponding speech recognition result can be obtained, including the speech recognition text "performing audio-to-text conversion while recording" and the timestamps of the sentences and words therein. For "achieving synchronization of speech and text", the corresponding speech recognition result can be obtained, including the speech recognition text "achieving synchronization of speech and text" and the timestamps of the sentences and words therein. For "automatically recording meeting information", the corresponding speech recognition result can be obtained, including the speech recognition text "automatically recording meeting information" and the timestamps of the sentences and words therein, etc.
[0056] Taking "achieving synchronization of speech and text" as an example, the timestamp of the sentence can be 2000, and the timestamps of each word can be: 2000, 2160, 2310, 2480, 2700, and 2850 in sequence. Additionally, in practical applications, the speech recognition result can further include the time offset of the end of the audio corresponding to the sentence relative to the establishment time of the long connection. For example, the time offset of the end of the audio corresponding to "performing audio-to-text conversion while recording" relative to the establishment time of the long connection is 1600, and the time offset of the end of the audio corresponding to "achieving synchronization of speech and text" relative to the establishment time of the long connection is 3000, etc.
[0057] In an embodiment of the present disclosure, for each client, after obtaining the speech recognition result returned by ASR each time, for each alignment object in the speech recognition text, the corresponding position of the alignment object in the meeting recording file can be determined respectively according to the difference between the size of the current meeting recording file and the theoretical size, the start time of the meeting, the establishment time of the long connection, and the timestamp of the alignment object, etc.
[0058] In an embodiment of the present disclosure, for each alignment object, the third difference between the establishment time of the long connection and the start time of the meeting can be obtained, and the second product of the third difference and a preset constant can be obtained. Additionally, the third product of the timestamp of the alignment object and the constant can be obtained. Furthermore, the sum of the second product, the third product, and the difference between the size of the current meeting recording file and the theoretical size can be obtained to get the file offset corresponding to the alignment object, and then the corresponding position of the alignment object in the meeting recording file can be determined according to the file offset.
[0059] That is, rec_offset = (asr_start_time – conf_start_time)*K + start_time (or word_offset)*K + pcm_offset; (2)
[0060] Among them, asr_start_time indicates the establishment time of the long connection, conf_start_time indicates the start time of the conference, K indicates a constant, start_time indicates the timestamp of the sentence, word_offset indicates the timestamp of the word, pcm_offset indicates the difference between the size of the current conference recording file and the theoretical size, and rec_offset indicates the file offset.
[0061] In the above formula, when the alignment object is a sentence, the third product between the timestamp of the alignment object and the constant is start_time*K, and when the alignment object is a word, the third product between the timestamp of the alignment object and the constant is word_offset*K. Taking the sentence "realizing the synchronization of sound and words" as an example, assuming that the timestamps of the words therein are 2000, 2160, 2310, 2480, 2700 and 2850, then substituting 2000 into the above formula (2) can obtain the file offset corresponding to the word "real", substituting 2160 into the above formula (2) can obtain the file offset corresponding to the word "present", substituting 2310 into the above formula (2) can obtain the file offset corresponding to the word "sound", and so on.
[0062] For any alignment object, after obtaining the file offset of the alignment object, the corresponding position of the alignment object in the conference recording file can be determined accordingly.
[0063] It can be seen that in the above processing method, with the help of the difference between the current size of the conference recording file and the theoretical size, the impact of the lag of the conference recording can be eliminated. Combined with the timestamp and other time information carried in the speech recognition results, the sentences and / or words in the speech recognition text can be accurately aligned to the corresponding positions in the conference recording file. In addition, not only can the sentence granularity alignment be achieved, but also the word granularity alignment can be achieved, that is, a finer-grained alignment can be achieved, thereby greatly improving the product experience. For example, the speech recognition text of the meeting can be searched by a certain keyword. After obtaining the search result, the corresponding position of this keyword in the conference recording file can be located with one click. For another example, the text in the corresponding speech recognition text can be displayed dynamically in real time when playing the conference recording, etc., thereby significantly improving the efficiency of meeting backtracking, etc.
[0064] In addition, in an embodiment of the present disclosure, when any client has an abnormality, the long connection corresponding to the client can be disconnected. When the client returns to normal, a long connection can be re-established for the client, and the establishment time of the re-established long connection can be used to update the previous establishment time.
[0065] During the progress of the meeting, situations such as the microphone of a certain client being turned off, dropping the line, leaving, etc. may occur, resulting in an abnormality of the client. In these situations, the long connection corresponding to the client can be disconnected. When the client returns to normal, such as after reconnecting to the meeting, a long connection can be re-established for the client, and the establishment time of the re-established long connection can be used to update the previous establishment time, that is, update asr_start_time. Correspondingly, when calculating the file offset of the alignment object corresponding to the client according to formula (2), the updated asr_start_time can be used, thus ensuring the accuracy of the calculation result, etc.
[0066] Based on the above introduction, Figure 5 is a flowchart of the second embodiment of the meeting information processing method described in the present disclosure. Assume that the clients in this embodiment include Client 1, Client 2, and Client 3, as Figure 5 shown, including the following specific implementation manners.
[0067] In step 501, when the meeting starts, record the start time of the meeting.
[0068] In step 502, establish long connections between Client 1, Client 2, and Client 3 that access the meeting and ASR respectively, and record the establishment time of the long connections.
[0069] In step 503, obtain the audio streams of Client 1, Client 2, and Client 3 respectively.
[0070] In step 504, decode the obtained audio stream into the PCM format.
[0071] In step 505, mix the audio streams of different clients into a single-channel audio and write it into the meeting recording file.
[0072] In step 506, for each of Client 1, Client 2, and Client 3, process it respectively in the manner described in steps 507 to 508.
[0073] In step 507, send the audio stream of the client to ASR through the long connection and obtain the speech recognition result returned by ASR.
[0074] The speech recognition result may include the speech recognition text and the timestamps of the alignment objects in the speech recognition text. The timestamp is the time offset of the start of the audio corresponding to the alignment object relative to the establishment time of the long connection. The alignment objects include at least one of the following: sentence, character, and word.
[0075] In step 508, for each obtained speech recognition result, respectively obtain the difference between the size of the current conference recording file and the theoretical size, and for each alignment object in the speech recognition text of the speech recognition result, respectively determine the corresponding position of the alignment object in the conference recording file according to the difference, the start time of the conference, the establishment time of the long connection, and the timestamp of the alignment object.
[0076] Among them, the first difference between the current time and the start time of the conference can be obtained, and the first product between the first difference and a preset constant can be obtained. The constant represents the file size of each millisecond of audio recording. Furthermore, the second difference between the first product and the size of the current conference recording file can be obtained, and the second difference is used as the difference between the size of the current conference recording file and the theoretical size.
[0077] In addition, for any alignment object, the third difference between the establishment time of the long connection corresponding to the alignment object and the start time of the conference can be obtained, and the second product between the third difference and the constant can be obtained. The third product between the timestamp of the alignment object and the constant can also be obtained, and the sum of the second product, the third product, and the difference can be obtained to obtain the file offset corresponding to the alignment object. Furthermore, the corresponding position of the alignment object in the conference recording file can be determined according to the file offset.
[0078] In step 509, the conference recording file, the speech recognition text, and the alignment information are stored in the cloud.
[0079] The alignment information refers to the corresponding position information of each alignment object in the conference recording file.
[0080] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present disclosure. In addition, for the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions in other embodiments.
[0081] The above is the introduction of the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.
[0082] Figure 6 This is a schematic structural diagram of Embodiment 600 of the conference information processing device described in the present disclosure. As Figure 6 shown, it includes: an acquisition module 601, a mixing module 602, and an alignment module 603.
[0083] The acquisition module 601 is used to acquire the audio stream of the client accessing the conference.
[0084] The mixing module 602 is used to mix the audio streams of different clients into a single-channel audio and write it into the conference recording file.
[0085] The alignment module 603 is used to perform the following processing for any client: obtain the speech recognition result of the audio stream of this client, and obtain the difference between the size of the conference recording file and the theoretical size, and align the speech recognition text in the speech recognition result with the conference recording file according to the difference.
[0086] Adopting the solution described in the above device embodiment, during the conference, a conference recording file can be generated according to the acquired audio stream of the client and the corresponding speech recognition result can be obtained. Moreover, for each speech recognition result obtained for each client each time, the difference between the size of the current conference recording file and the theoretical size can be obtained, and the speech recognition text in the speech recognition result can be aligned with the conference recording file according to the difference, that is, the lag of the conference recording is considered during alignment, thereby improving the accuracy of the alignment result, etc.
[0087] In the solution described in the present disclosure, after the conference starts, as each client accesses, the acquisition module 601 can acquire the audio streams of each client, and the mixing module 602 can mix the audio streams of different clients into a single-channel audio and write it into the conference recording file.
[0088] Due to reasons such as the network, the audio streams at the same moment on different clients cannot arrive completely simultaneously. Therefore, the mixing algorithm will perform a certain caching process on each audio stream, and a mixing will be performed only when a certain condition is met. The specific condition depends on different mixing algorithms. Correspondingly, there will be a certain lag in the conference recording, which also brings difficulties to the alignment of the conference recording file and the speech recognition text.
[0089] To solve the above problems, in the solution described in the present disclosure, a variable pcm_size can be used to record the size of the current conference recording file. Through the size of the current conference recording file and the start time of the conference, etc., the difference between the size of the current conference recording file and the theoretical size can be determined.
[0090] Accordingly, at the beginning of the meeting, the start time of the meeting can be recorded. In one embodiment of the present disclosure, the alignment module 603 can obtain a first difference between the current time and the start time of the meeting, and can obtain a first product between the first difference and a preset constant, where the constant represents the file size of audio recording per millisecond. Furthermore, a second difference between the first product and the size of the current meeting recording file can be obtained, and the second difference is used as the difference between the size of the current meeting recording file and the theoretical size.
[0091] For each client, the alignment module 603 can also perform the following processing respectively: obtain the speech recognition result of the audio stream of the client, and obtain the difference between the size of the current meeting recording file and the theoretical size, and align the speech recognition text in the speech recognition result with the meeting recording file according to the difference.
[0092] In one embodiment of the present disclosure, for each client, the alignment module 603 can respectively establish a long connection with the ASR for the client, and can send the audio stream of the client to the ASR through the long connection, and obtain the speech recognition result returned by the ASR. Preferably, the long connection can be a websocket long connection.
[0093] For each client, the establishment time of the long connection can also be recorded respectively. During the process of continuously pushing the audio stream to the ASR, the ASR will continuously return the speech recognition result through the websocket callback.
[0094] In one embodiment of the present disclosure, in addition to the speech recognition text, the speech recognition result can further include the time stamp of the alignment object in the speech recognition text, where the time stamp is the time offset of the start of the audio corresponding to the alignment object relative to the establishment time of the long connection, and the alignment object includes at least one of the following: sentence, character, word.
[0095] Accordingly, in one embodiment of the present disclosure, for each client, after the alignment module 603 obtains the speech recognition result returned by the ASR each time, for each alignment object in the speech recognition text, it can respectively determine the corresponding position of the alignment object in the meeting recording file according to the difference between the size of the current meeting recording file and the theoretical size, the start time of the meeting, the establishment time of the long connection, and the time stamp of the alignment object, etc.
[0096] In one embodiment of the present disclosure, for each alignment object, the alignment module 603 may obtain a third difference between the establishment time of the long connection corresponding to the alignment object and the start time of the meeting, and may obtain a second product of the third difference and a preset constant. Additionally, it may also obtain a third product of the timestamp of the alignment object and the constant. Furthermore, it may obtain the sum of the second product, the third product, and the difference between the size of the current meeting recording file and the theoretical size, to obtain the file offset corresponding to the alignment object. Then, it may determine the corresponding position of the alignment object in the meeting recording file according to the file offset.
[0097] In addition, in one embodiment of the present disclosure, when any client has an abnormality, the alignment module 603 may disconnect the long connection corresponding to the client. When the client returns to normal, it may re - establish a long connection for the client and may update the previous establishment time with the establishment time of the re - established long connection.
[0098] Figure 6 The specific working process of the illustrated device embodiment may refer to the relevant descriptions in the foregoing method embodiment.
[0099] The solution described in the present disclosure may be applied to the field of artificial intelligence, particularly in the fields of intelligent speech and natural language processing, etc. Artificial intelligence is a discipline that studies how to make a computer simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware - level technologies and software - level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0100] The voice in the embodiments described in the present disclosure is not the voice of a specific user and does not reflect the personal information of a specific user. Additionally, the execution subject of the meeting information processing method may obtain the voice through various public, legal, and compliant means, such as obtaining it from the user after the user's authorization. In short, in the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0101] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0102] Figure 7FIG. 0 shows a schematic block diagram of an electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0103] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0104] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0105] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the methods described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described in this disclosure can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the methods described in this disclosure by any other suitable means (e.g., by means of firmware).
[0106] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0107] The program code for implementing the methods of this disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0108] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0110] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0111] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0112] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0113] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for processing conference information, comprising: Obtain the audio stream of the client accessing the conference; Mix the audio streams of different clients into a single-channel audio and write it into the conference recording file; For any client, perform the following processing respectively: Obtain the speech recognition result of the audio stream of the client, where the speech recognition result includes the speech recognition text and the time stamp of the alignment object in the speech recognition text, and the time stamp is the time offset of the start of the audio corresponding to the alignment object relative to the establishment time of the long connection. The alignment object includes at least one of the following: sentence, character, word. The long connection is the long connection between the client and the automatic speech recognition system; Obtain the difference between the size of the conference recording file and the theoretical size; For any alignment object, respectively determine the corresponding position of the alignment object in the conference recording file according to the difference, the start time of the conference, the establishment time of the long connection, and the time stamp of the alignment object.
2. The method according to claim 1, wherein, The obtaining the speech recognition result of the audio stream of the client includes: Establish the long connection for the client; Send the audio stream of the client to the automatic speech recognition system through the long connection and obtain the speech recognition result returned by the automatic speech recognition system.
3. The method according to claim 1 or 2, wherein, The obtaining the difference between the size of the conference recording file and the theoretical size includes: Obtain the first difference between the current time and the start time of the conference; Obtain the first product of the first difference and a preset constant, where the constant represents the file size of audio recording per millisecond; Obtain the second difference between the first product and the size of the conference recording file, and use the second difference as the difference between the size of the conference recording file and the theoretical size.
4. The method according to claim 3, wherein, The respectively determining the corresponding position of the alignment object in the conference recording file according to the difference, the start time of the conference, the establishment time of the long connection, and the time stamp of the alignment object includes: Obtain the third difference between the establishment time of the long connection and the start time of the conference; Obtain the second product of the third difference and the constant; Obtain the third product of the time stamp of the alignment object and the constant; Obtain the sum of the second product, the third product, and the difference to obtain the file offset corresponding to the alignment object; Determine the corresponding position of the alignment object in the conference recording file according to the file offset.
5. The method according to claim 2, further comprising: When any client has an abnormality, disconnect the long connection corresponding to the client. When the client returns to normal, re-establish the long connection for the client and update the previous establishment time with the establishment time of the re-established long connection.
6. A conference information processing apparatus, comprising: An obtaining module, a mixing module, and an alignment module; The obtaining module is used to obtain the audio stream of the client accessing the conference; The mixing module is used to mix the audio streams of different clients into a single-channel audio and write it into the conference recording file; The alignment module is configured to perform the following processing for any client: obtain the speech recognition result of the audio stream of the client, where the speech recognition result includes the speech recognition text and the timestamps of the alignment objects in the speech recognition text, the timestamps being the time offsets of the start of the audio corresponding to the alignment objects relative to the establishment time of the long connection, the alignment objects including at least one of the following: sentences, characters, words, and the long connection being the long connection between the client and the automatic speech recognition system; obtain the difference between the size of the conference recording file and the theoretical size; for any alignment object, respectively determine the corresponding position of the alignment object in the conference recording file according to the difference, the start time of the conference, the establishment time of the long connection, and the timestamp of the alignment object.
7. The apparatus according to claim 6, wherein, The alignment module establishes the long connection for the client, sends the audio stream of the client to the automatic speech recognition system through the long connection, and obtains the speech recognition result returned by the automatic speech recognition system.
8. The apparatus according to claim 6 or 7, wherein, The alignment module obtains a first difference between the current time and the start time of the conference, obtains a first product of the first difference and a preset constant, where the constant represents the file size of audio recording per millisecond, obtains a second difference between the first product and the size of the conference recording file, and uses the second difference as the difference between the size of the conference recording file and the theoretical size.
9. The apparatus according to claim 8, wherein, The alignment module obtains a third difference between the establishment time of the long connection and the start time of the conference, obtains a second product of the third difference and the constant, and obtains a third product of the timestamp of the alignment object and the constant, obtains the sum of the second product, the third product, and the difference, to obtain the file offset corresponding to the alignment object, and determines the corresponding position of the alignment object in the conference recording file according to the file offset.
10. The apparatus according to claim 7, wherein, The alignment module is further configured to, when any client has an abnormality, disconnect the long connection corresponding to the client, and when the client returns to normal, re-establish the long connection for the client, and update the previous establishment time with the establishment time of the re-established long connection.
11. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-5.
13. A computer program product comprising a computer program / instructions which, when executed by a processor, implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Multimedia transliteration method and system
CN105895085A
Conference voice data processing method and device, computer equipment and storage medium
CN110322872A