Epidemic investigation information processing method and device, storage medium, and electronic device
Through the external device, the audio streams of interviewers and target objects are merged and combined with template information, the flow investigation information is automatically generated, which solves the problem of low efficiency and poor timeliness generation of flow investigation information in the existing technology, and realizes efficient and convenient flow investigation information processing.
Patent Information
- Application Number
- CN202210426728.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-21
AI Technical Summary
In the prior art, the generation efficiency of the flow investigation information is low and the timeliness is poor, and it mainly relies on cumbersome manual operations.
The audio stream in the conversation request between the interviewer and the target object is obtained through the external device, the first input audio stream and the second input audio stream are merged, structured processing is performed, and the streaming investigation information is generated based on the template information.
It realizes the automatic generation of flow investigation information, improves the generation efficiency and timeliness, and enhances the portability and flexibility of flow investigation operations.
Smart Images

Figure CN114822549B_ABST
Abstract
Description
Background Art
[0002] After an infectious disease occurs, epidemiological investigations of infected cases and their associated cases are of extremely important significance for the prevention and control of infectious diseases.
[0003] In related technologies, generally, epidemiological investigators mainly have conversations with infected cases and their associated cases. During the conversation, characteristic information is manually extracted, and the characteristic information is manually sorted into an epidemiological investigation report according to the characteristic information. In the above method, since it is completed through manual operations, the operation steps are relatively cumbersome, the operation efficiency is low, and the timeliness is poor.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present disclosure provides a method and apparatus for processing epidemiological investigation information, a computer-readable storage medium, and an electronic device, thereby at least to some extent overcoming the problem of low generation efficiency of epidemiological investigation information in related technologies.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, there is provided a method for processing epidemiological investigation information, including: responding to a conversation request of an interviewer and a target object sent by a sending end, and obtaining a first input audio stream in the conversation request based on an external device; obtaining a second input audio stream of the interviewer; merging the first input audio stream and the second input audio stream to obtain a mixed audio stream; performing structured processing on the mixed audio stream to determine characteristic information, and determining the epidemiological investigation information of the target object in combination with template information and the characteristic information.
[0008] In an exemplary embodiment of the present disclosure, the obtaining the first input audio stream in the conversation request based on the external device includes: obtaining audio streams of a plurality of audio input sources, where the audio streams include a first type identifier and a second type identifier, and wherein the first type identifier is used to determine the second input audio stream, and the second type identifier is used to determine the first input audio stream; screening the audio streams of the plurality of audio input sources through the first type identifier to determine the second input audio stream from the audio streams; determining the value of the second type identifier of the second input audio stream, and determining the audio stream different from the value of the second type identifier as the first input audio stream of the external device.
[0009] In an exemplary embodiment of the present disclosure, before merging the first input audio stream and the second input audio stream to obtain a mixed audio stream, the method further includes: determining, according to the second type identifier, the same audio streams in the first input audio stream and / or the second input audio stream as duplicate audio streams, and performing a filtering operation on the duplicate audio streams.
[0010] In an exemplary embodiment of the present disclosure, the merging of the first input audio stream and the second input audio stream to obtain a mixed audio stream includes: merging the first input audio stream and the second input audio stream in chronological order to obtain the mixed audio stream.
[0011] In an exemplary embodiment of the present disclosure, before merging the first input audio stream and the second input audio stream to obtain a mixed audio stream, the method further includes: performing a sound effect adjustment operation on the real-time audio parameters that meet the audio conditions in the first input audio stream and the second input audio stream to adjust the first input audio stream and the second input audio stream.
[0012] In an exemplary embodiment of the present disclosure, the structuring the mixed audio stream to determine feature information, and combining the template information and the feature information to determine the epidemiological investigation information of the target object includes: obtaining audio data corresponding to the mixed audio stream;
[0013] Converting the audio data into text information, extracting feature information from the text information, and filling the feature information with the template information to generate the epidemiological investigation information.
[0014] In an exemplary embodiment of the present disclosure, the obtaining audio data corresponding to the mixed audio stream includes: if the scenario corresponding to the dialogue request is a first type of scenario, performing a recording operation on the mixed audio stream according to the recording parameters to obtain audio data that matches the speech recognition requirements; the audio parameters include one or a combination of a recording format, a sampling rate, a sampling bit depth, and the number of channels; if the scenario corresponding to the dialogue request is a second type of scenario, using the mixed audio stream as the audio data.
[0015] According to one aspect of the present disclosure, there is provided an epidemiological investigation information processing device, including: a first audio stream acquisition module, configured to respond to a conversation request between an interviewer and a target object sent by a sending end, and acquire a first input audio stream in the conversation request based on an external device; a second audio stream acquisition module, configured to acquire a second input audio stream of the interviewer; an audio merging module, configured to merge the first input audio stream and the second input audio stream to obtain a mixed audio stream; and an epidemiological investigation information generation module, configured to perform structured processing on the mixed audio stream to determine feature information, and determine the epidemiological investigation information of the target object in combination with template information and the feature information.
[0016] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the epidemiological investigation information processing method described in any one of the above is implemented.
[0017] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein the processor is configured to execute the epidemiological investigation information processing method described in any one of the above by executing the executable instructions.
[0018] In the embodiments of the present disclosure, the epidemiological investigation information processing method, the epidemiological investigation information processing device, the computer-readable storage medium, and the electronic device acquire a corresponding first input audio stream from the information of the conversation request associated with the target object based on an external device, merge the first input audio stream and the second input audio stream of the interviewer into a mixed audio stream, and then determine the epidemiological investigation information in combination with the template information and the feature information corresponding to the mixed audio stream. On the one hand, the receiving end receives the first input audio stream in the conversation request between the interviewer and the target object sent by the external device and receives the second input audio stream of the interviewer, and then generates a mixed audio stream, and automatically generates epidemiological investigation information according to the template information and the feature information of the mixed audio stream, realizing the process of automatically generating epidemiological investigation information, reducing the operation steps in manual operation, improving the generation efficiency of epidemiological investigation information, and improving timeliness. On the other hand, the first input audio stream can be extracted from the conversation request based on the external device, the receiving end receives the first input audio stream of the conversation request sent by the external device and receives the second input audio stream input by the built-in microphone, and then generates epidemiological investigation information according to the mixed audio stream generated by the first input audio stream and the second input audio stream. Since the external device can be installed on the receiving end, the accessories are simple and easy to configure, improving the portability and flexibility of the epidemiological investigation operation, and being capable of large-scale promotion, increasing the application scope and executability.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0021] Figure 1 Schematically shows the system architecture diagram for implementing the epidemic investigation information processing method in an embodiment of the present disclosure.
[0022] Figure 2 Schematically shows a schematic diagram of an epidemic investigation information processing method in an embodiment of the present disclosure.
[0023] Figure 3 Schematically shows the architecture diagram for audio stream processing in an embodiment of the present disclosure.
[0024] Figure 4 Schematically shows the process diagram for audio stream merging in an embodiment of the present disclosure.
[0025] Figure 5 Schematically shows the process diagram for generating epidemic investigation information in an embodiment of the present disclosure.
[0026] Figure 6 Schematically shows the process diagram for telephone epidemic investigation in an embodiment of the present disclosure.
[0027] Figure 7 Schematically shows the block diagram of the epidemic investigation information processing device in an embodiment of the present disclosure.
[0028] Figure 8 Schematically shows the block diagram of an electronic device in an embodiment of the present disclosure. Detailed implementation manners
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0030] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0031] An epidemiological investigation information processing method is provided in an embodiment of the present disclosure. Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present disclosure can be applied is shown.
[0032] As Figure 1 shown, the system architecture 100 may include a sending end 101, an external device 102, and a receiving end 103. Among them, the sending end may be a terminal device capable of making a call or having a conversation, such as a smart phone, a tablet computer, a smart watch, a smart bracelet, a smart speaker, etc. The external device 102 is used to communicatively connect the sending end 101 and the receiving end 103. The external device is used to obtain the conversation request between the interviewer and the target object sent by the sending end, use the information in the conversation request as its first input audio stream, and send the first input audio stream to the receiving end. In the embodiment of the present disclosure, the external device 102 between the sending end 101 and the receiving end 103 may be a USB external sound card. The sending end 101 and the external device 102 are connected by an audio cable. The receiving end 103 may be a terminal device with computing functions, such as a portable computer, a desktop computer, a smart phone, etc. with computing functions, and is used to process the data sent by the sending end.
[0033] In the embodiments of the present disclosure, the sending end 101 is used to obtain the information in the conversation request, and the external device 102 is installed on the receiving end. The external device is used to transmit the information in the sending end to the receiving end. The external device is connected to the sending end through an audio cable, and after the connection is completed, a communication connection between the sending end and the receiving end is established through the external device, and the information of the conversation request of the sending end is imported into the input end of the receiving end. The receiving end 103 is used to receive the first input audio stream of the conversation request sent by the external device 102 and the second input audio stream input by the built-in microphone, and merge the first input audio stream and the second input audio stream to obtain a mixed audio stream. Further, the receiving end 103 also performs structured processing on the mixed audio stream to obtain feature information, and automatically generates the epidemiological investigation information corresponding to the target object in combination with the template information and the feature information.
[0034] It should be noted that the epidemiological investigation information processing method provided by the embodiments of the present disclosure can be executed by the receiving end, and can be specifically implemented according to the computer program stored on the receiving end.
[0035] Based on the above system architecture, in the embodiments of the present disclosure, an epidemiological investigation information processing method is provided, which is applied to the receiving end and is used to realize automatic epidemiological investigation through an external device and the receiving end in a call scenario. Refer to Figure 2 As shown in, this epidemiological investigation information processing method includes steps S210 to S230, which are introduced in detail as follows:
[0036] In step S210, in response to the conversation request of the interviewer and the target object sent by the sending end, the first input audio stream in the conversation request is obtained based on the external device.
[0037] In the embodiments of the present disclosure, the target object can be an object associated with the application scenario, and can be specifically different according to different application scenarios. When the application scenario is an epidemiological investigation scenario of an infectious disease, the target object can be a person associated with the infectious disease, such as a confirmed case of an infectious disease or a person who has come into contact with an infectious disease, etc. When the application scenario is a questionnaire survey or a telephone return visit, the target object can be a user associated with the questionnaire survey and the telephone return visit, such as a user who uses a certain product. Here, the application scenario is taken as an example of an epidemiological investigation scenario of an infectious disease for illustration. The infectious disease can be a contagious disease (infectious disease), such as various types of epidemics or various contagious influenza, etc. The infectious disease can be for a certain region or for all regions, and is not limited here. The unit time can be, for example, every day, or every two days or every week, etc.
[0038] The dialogue request can be established based on the sending end. The sending end can enable the interviewer and the target object to establish a dialogue request through methods such as video calls, phone calls, voice conversations, instant messaging application conversations, etc. For example, if it is detected that the terminal 1 (sending end) of the interviewer communicates with the terminal 2 of the target object through a phone call, it can be considered that a dialogue request between the interviewer and the target object established by the sending end is detected. During the dialogue process, the sending end can play and output the information of the dialogue request. The information of the dialogue request can include any type of information associated with the target object, such as the name, phone number, age, identity information, location information, historical track information, and protection information (whether vaccinated, whether wearing a mask, whether wearing protective items) of the target object, etc.
[0039] After detecting the dialogue request, the external device can import the information of the dialogue request into the receiving end. The external device can be an external sound card, that is, a sound card connected to the motherboard through an interface. For example, it can be a USB external sound card directly plugged into the USB interface. The first end of the external sound card is connected to the audio cable and is used to connect to the sending end through the audio cable; the second end of the external sound card is connected to the receiving end and is used to import the information of the dialogue request of the sending end into the receiving end. As Figure 3 shown in the architecture diagram, the USB external sound card 302 is connected to the computer (receiving end) 303. The mobile phone (sending end) 301 is connected to the mobile phone audio cable 304, and the output end of the mobile phone audio cable 304 is inserted into the MIC input interface of the external sound card. Among them, the external sound card cannot use the earphone-microphone integrated interface and needs to use an independent microphone input interface.
[0040] After establishing a communication connection between the sending end and the receiving end through the external device, the first input audio stream in the dialogue request can be transmitted to the input end of the receiving end as the input information of the receiving end through the external device. This process can be implemented in multiple programming languages. Here, the JavaScript language on the web page is used as an example for illustration.
[0041] In some embodiments, the first input audio stream and the second input audio stream obtained based on an external device can be determined through identification information. Among them, the first input audio stream and the second input audio stream can be determined through a first type identifier and a second type identifier. Specifically, audio streams of multiple audio input sources can be obtained through an interface, and the audio streams of the multiple audio input sources can be filtered through the first type identifier to determine the second input audio stream from the audio streams; after distinguishing the second input audio stream, the second type identifier of the second input audio stream can be obtained, the value of the second type identifier of the second input audio stream can be determined, and the audio stream different from the value of the second type identifier can be determined as the first input audio stream obtained by the external device. Among them, each audio stream can include a first type identifier and a second type identifier. The first type identifier can be an input device source identifier deviceId, which is used to identify whether the audio stream belongs to the built-in microphone input, that is, the first type identifier is used to determine whether the audio stream belongs to the second input audio stream. The identification meaning of the second type identifier is different from that of the first type identifier. For example, the second type identifier can be groupId, and groupId represents a group identifier, which is used to distinguish whether the audio stream is input by the built-in microphone itself or by an external device, that is, the second type identifier is used to determine whether the audio stream belongs to the first input audio stream.
[0042] Through the interface Web Audio API provided by the browser, all audio input sources can be obtained. If the first type identifier of an audio input source, that is, the input device source identifier deviceId, is the target value default, it is determined that the audio input source is the built-in MIC input source of the receiving end, and the audio stream of this audio input source can be considered as the second input audio stream. After determining the built-in MIC input of the receiving end, the audio streams of multiple audio input sources can be further filtered based on the second type identifier to determine the first input audio stream obtained by the external device. Exemplarily, the audio stream different from the value of the above-mentioned second type identifier among the audio streams of multiple audio input sources can be determined as the first input audio stream of the external device. When the second input audio stream is determined, the value of the second type identifier of the second input audio stream is further determined; when there is an audio stream different from the value of the second type identifier, it is considered that this audio stream is the first input audio stream obtained by the external device. That is, the audio stream with a second type identifier different from that of the built-in microphone input of the receiving end is determined as the first input audio stream sent by the external device. In some embodiments, by comparing the second type identifier groupId, the audio input with a second type identifier groupId different from the groupId value of the built-in MIC input of the receiving end is determined as the audio input of the external device. Through the first type identifier and the second type identifier, the type and source of the audio stream can be accurately distinguished.
[0043] In step S220, obtain the second input audio stream of the interviewer.
[0044] In the embodiments of the present disclosure, the interviewer can be an epidemiological investigator, or a smart assistant or a smart robot, etc., as long as it can conduct a call-based epidemiological investigation on the target object. The second input audio stream can be input through the built-in microphone of the receiving end, that is, the second input audio stream can be the audio stream of the interviewer himself / herself, and during the conversation, the second input audio stream can be directly stored in the receiving end. The second input audio stream can be an audio stream related to the conversation request, or an audio stream unrelated to the conversation request, which is not limited here, as long as it is the audio stream of the interviewer during the conversation request process.
[0045] For example, if interviewer A sends a conversation request to target object B, the first input audio stream can be the audio stream in the conversation request sent by an external device, which can include the audio streams of interviewer A and target object B; the second input audio stream can be the audio stream input by interviewer A through the built-in microphone of the receiving end.
[0046] On this basis, the first input audio stream in the conversation request can be obtained through an external device, and the first input audio stream is sent to the receiving end through the external device. The receiving end is used to receive the first input audio stream of the conversation request sent by the external device and the second input audio stream input by the built-in microphone. After the communication connection between the sending end and the receiving end is established through the external device, the audio output on the mobile phone (sending end) can be played and output through the speaker or earphone on the computer (receiving end).
[0047] Refer to Figure 3 As shown in, the USB external sound card 302 is connected to the computer 303. The mobile phone 301 is communicatively connected to the external sound card through the mobile phone audio cable 304, and the output end of the mobile phone audio cable 304 is inserted into the MIC input interface of the external sound card, thereby realizing the communication connection between the sending end and the receiving end. During the process of the epidemiological investigator (interviewer) making a conversation request with the target object through the mobile phone, the information of the conversation request flows out through the mobile phone audio cable, and then is sent to the built-in speaker or the headphone audio output interface in the computer through the external sound card for output, so that the computer receives the first input audio stream in the conversation request sent by the external device. At the same time, the voice of the epidemiological investigator (interviewer) is input into the computer through the built-in MIC (microphone) interface of the computer, so that the computer receives the second input audio stream.
[0048] Continue to refer to Figure 2 As shown in, in step S230, merge the first input audio stream and the second input audio stream to obtain a mixed audio stream.
[0049] In the embodiments of the present disclosure, after obtaining the first input audio stream from the conversation request of the sending end through an external device and importing the first input audio stream to the receiving end through the external device, the receiving end may merge the first input audio stream in the conversation request sent by the external device and the second input audio stream input by the built-in microphone of the receiving end to obtain a mixed audio stream.
[0050] To improve the accuracy of the audio stream, duplicate audio streams in the first input audio stream and / or the second input audio stream may be filtered according to the second type identifier. The same audio stream in the first input audio stream or the second input audio stream may be determined as a duplicate audio stream, and the same audio stream in the first input audio stream and the second input audio stream may also be determined as a duplicate audio stream, and the duplicate audio stream may be filtered. That is, the audio streams of all audio input sources are obtained, and a duplicate filtering operation is performed through the groupId. By performing a duplicate filtering operation on the first input audio stream and / or the second input audio stream according to the second type identifier, the problem of repeated merging caused by duplicate audio streams can be avoided, the accuracy of merging can be improved, and resource waste can be avoided. Performing a duplicate filtering operation according to the second type identifier means filtering the first input audio stream corresponding to the second type identifier or the second input audio stream respectively, and filtering the first input audio stream and the second input audio stream to delete the duplicate audio streams contained therein. By deleting the duplicate audio streams in all audio sources, the interference of duplicate audio streams can be avoided, and the problem of low merging accuracy caused by duplicate audio streams can be reduced. Exemplarily, the content of the first input audio stream and the second input audio stream is compared. If the content is the same, the partial audio stream corresponding to the same content in the first input audio stream and the second input audio stream is determined as a duplicate audio stream; if the content is different, the first input audio stream and the second input audio stream are determined as non-duplicate audio streams. For example, if a part of the content 1 of the audio stream a in the first input audio stream is the same as the audio stream b in the second input audio stream, then the part of the content 1 of the two is considered to be a duplicate audio stream. On this basis, the duplicate audio stream can be filtered to improve the accuracy.
[0051] When merging the audio streams of different audio input sources, the first input audio stream and the second input audio stream can also be adjusted in sound effect and noise reduction according to real-time audio parameters to optimize the first input audio stream and the second input audio stream. The real-time audio parameters can include, but are not limited to, one or a combination of more of volume, echo, noise, and timbre. The sound effect adjustment operation can be performed according to the real-time audio parameters. For example, it can include, but is not limited to, volume adjustment, echo processing, noise cancellation, timbre modification, etc. When the real-time audio parameters meet the audio conditions, the sound effect adjustment operation can be performed on the real-time audio parameters that meet the audio conditions. Among them, any one or more of the situations where the volume is greater than the first preset value or less than the second preset value, there is an echo, and the noise is greater than the third preset value can be considered that the real-time audio parameters meet the audio conditions. Among them, the first preset value can be greater than the second preset value, and the third preset value can be greater than the first preset value. Or it can also be determined whether to perform the sound effect adjustment operation on the real-time audio parameters according to the actual application scenario. The operation degree of the sound effect adjustment operation can be determined according to the magnitude of the real-time audio parameters and the reference value in the audio conditions. The reference value can be the first preset value, the second preset value, or the third preset value. In addition, the operation degree of the sound effect adjustment operation can also be determined according to other parameters, as long as the processed real-time audio parameters meet the audio conditions. For example, the sound effect adjustment operation can be implemented on the first input audio stream and the second input audio stream, such as increasing or decreasing the volume, reducing the echo, reducing the noise, adjusting the timbre, etc. If the volume is greater than the first preset value, the volume can be turned down so that it is less than the first preset value.
[0052] Next, the first input audio stream and the second input audio stream can be merged in chronological order. Merging in chronological order can include: merging the audio streams from different audio input sources within the same time period, that is, splicing and merging the first input audio stream of the external device and the second input audio stream of the receiving end within the same time period. Specifically, it can be merged according to the start times of the first input audio stream and the second input audio stream. If the start times of the first input audio stream and the second input audio stream are the same within the same time period, the two can be mixed. If the start times of the first input audio stream and the second input audio stream are different within the same time period, the first input audio stream and the second input audio stream are spliced and merged according to the arrangement order of the start times. The arrangement order of the start times refers to the order of the start times. Merging in chronological order can also include: synthesizing the audio streams from different audio input sources in the order of time. For example, if the time of the first input audio stream a is later than the time of the second input audio stream b, the first input audio stream and the second input audio stream can be synthesized into a mixed audio stream composed of the second input audio stream b and the first input audio stream a.
[0053] In some embodiments, the first input audio stream of the conversation request sent by the USB external sound card received by the receiving end and the audio stream input of the built-in microphone of the receiving end can be merged. This can be achieved through multiple programming languages. Here, the JavaScript language on the web page is used as an example. Through the audio merging method in the Web Audio API interface provided by the browser, all audio input sources are obtained and filtered once through the second type identifier groupId. Further, an audio processing flow is established using the AudioContext class, and multiple audio sources are connected to the same audio destination to achieve aggregation. The AudioContext class represents an audio processing graph constructed by connected audio modules and is used to build a pipeline for multiple different sources to merge the audio streams of multiple different sources. It should be added that multiple processing nodes can be embedded in the AudioContext class to implement various sound effect adjustment operations, and each node is used to execute one sound effect adjustment operation.
[0054] In some embodiments, since the audio stream can be represented as waveform data, that is, the function corresponding to the audio stream can be Fourier-expanded to generate sine waves and cosine waves. Therefore, each sine wave and cosine wave can be modified by modifying parameters such as its corresponding amplitude, phase, wavelength, etc., and the modified sine waves and cosine waves are merged in chronological order to obtain a mixed audio stream.
[0055] Reference Figure 4 As shown in [], the first input audio stream 401 and the second input audio stream 402 are respectively subjected to sound effect adjustment operations to obtain the processed first input audio stream 403 and the processed second input audio stream 404, and the processed first input audio stream 403 and the processed second input audio stream 404 are merged to obtain a mixed audio stream 405.
[0056] In the embodiments of the present disclosure, by merging the first input audio stream of the conversation request sent by the external device received by the receiving end and the second input audio stream of the built-in microphone received by the receiving end, the problem of incompleteness caused by missing audio streams is avoided, and accurate audio integration can be achieved, improving the integrity and comprehensiveness of the audio stream.
[0057] Continue to refer to Figure 2 As shown in [], in step S240, the mixed audio stream is structurally processed to determine feature information, and the flow adjustment information of the target object is determined by combining the template information and the feature information.
[0058] In the embodiments of the present disclosure, the structured processing is used to extract feature information to generate recognizable formatted text. The feature information may be keywords in the mixed audio stream, and the feature information may have an association relationship with the reference feature information in the template information. The reference feature information may be information used to represent name, age, ID number, address, contact status, historical trajectory information (time and historical location), and protection information. For example, if the reference feature information is name, the feature information in the mixed audio stream may be Zhang San, etc.; if the reference feature information is address, the feature information may be Community A, etc.
[0059] Figure 5 schematically shows a flowchart for generating epidemiological investigation information. Refer to Figure 5 as shown in, it mainly includes step S510 and step S520, where:
[0060] In step S510, audio data corresponding to the mixed audio stream is obtained.
[0061] In this step, the audio data can be used for feature information extraction. The audio data may be the mixed audio stream itself, or the data obtained by sampling the mixed audio stream. When the audio data is the mixed audio stream itself, the audio data is the audio data stream. When the audio data is the data obtained by sampling the mixed audio stream, the audio data is an audio file.
[0062] Such as Figure 5As shown in [description], obtaining audio data can include two methods: In step S511, if the scenario corresponding to the dialogue request is a first-type scenario, record the mixed audio stream according to the recording parameters to obtain audio data that matches the speech recognition requirements for storage. Among them, recording the mixed audio stream for storage facilitates subsequent playback or verification to timely update incorrect data. In some embodiments, the mixed audio stream can be recorded based on the recording parameters. The recording parameters can include, but are not limited to, one or a combination of recording format, sampling rate, sampling bit depth, and number of channels. Specifically, the MediaRecorder class can be used to record the mixed audio stream, with the recording format being PCM (Pulse Code Modulation) raw audio, the sampling rate being 16,000, the sampling bit depth being 16 bits, and the number of channels being 1. After the recording operation, an audio file that meets the speech recognition requirements can be obtained. Among them, the PCM format can convert analog signals such as sound into symbolized pulse trains for recording and storage. It should be noted that if the number of channels is more than 1, the redundant channel data is discarded. Among them, when the number of channels is greater than 1, it can be considered as stereo audio. Since computers generally use microphone arrays and the number of channels in the microphone array is relatively large. To support speech recognition, redundant channel data needs to be discarded to improve accuracy. The first-type scenario can be a scenario where the dialogue request meets the dialogue conditions, or it can be any type of scenario. If there is noise in the dialogue request, or the speech rate in the dialogue request is greater than the standard speech rate, or the amount of information in the dialogue request is large (e.g., greater than the information volume threshold), then it can be considered that the dialogue request meets the dialogue conditions. That is, the first-type scenario can be a scenario with noise, a speech rate greater than the standard speech rate, or an information volume greater than the information volume threshold.
[0063] After converting the mixed audio stream into audio data, the audio data can be input into a real-time speech recognition system for speech recognition.
[0064] In step S512, if the scenario corresponding to the dialogue request is a second-type scenario, use the mixed audio stream as the audio data. The second-type scenario can be a scenario without noise, a slow speech rate, or a small amount of information. If there is no noise in the dialogue request, or the speech rate in the dialogue request is not greater than the standard speech rate, or the amount of information in the dialogue request is small (e.g., not greater than the information volume threshold), then it can be considered that the dialogue request does not meet the dialogue conditions to determine the second-type scenario. That is, in the second-type scenario, the stream data corresponding to the mixed audio stream is directly transmitted in real time to connect the mixed audio stream to the real-time speech recognition system for speech recognition without the need to perform a recording operation to generate an audio file.
[0065] It should be added that the mixed audio stream can be converted into audio data in the first type of scenario, and the mixed audio stream can be directly used as audio data in the second type of scenario. In the embodiments of the present disclosure, taking the example that recording operations can be performed in any type of scenario and real-time transmission can be performed in any type of scenario, based on this, in any type of scenario, the mixed audio stream can be recorded according to audio parameters to obtain audio data, and the mixed audio stream can be directly used as audio data. Further, the audio data obtained by any one of the methods can be input into a real-time speech recognition system for speech recognition. Among them, one method can be selected according to the scenario type or actual requirements, which is not limited here. For example Figure 3 As shown in Figure 3 , the receiving end can perform audio synthesis on the first input audio stream of the conversation request sent by the external sound card received and the second input audio stream input by the built-in microphone received by the receiving end, and then perform format conversion to generate a target data stream, that is, audio data.
[0066] In step S520, the audio data is converted into text information, feature information is extracted from the text information, and the feature information is filled with template information to generate the epidemiological investigation information.
[0067] In this step, after obtaining the audio data, the audio data can be converted into text information by means of speech recognition. Further, the text information can be structurally processed to obtain feature information, that is, keywords. In some embodiments, the structural processing can be based on regular expressions or based on deep learning to extract feature information. Regular expressions refer to rule strings composed of characters and combinations of characters, which are used to filter strings. The deep learning-based method can be based on machine learning models, deep learning models, and other models that can be used to extract text information for extraction. For example, all analysis statements of the text information can be tokenized to obtain word units of the analysis statements; the word features of the word units, the statement features of the word units in the corresponding analysis statements, and the text features of the word units in the extracted text are obtained; based on a machine learning model established by a machine learning algorithm, the word features, statement features, and text features of the word units in each analysis statement are used to perform keyword extraction operations on each analysis statement. The machine learning model can be any type of model, such as a support vector machine, etc., which is not specifically limited here. In addition, for text information represented by word vectors, the words in the text information can be clustered by the K-Means algorithm, and the cluster center is selected as a main keyword of the text, and the distance between other words and the cluster center, that is, the similarity, is calculated, and the word closest to the cluster center is selected as the keyword. It should be noted that the speech recognition service in this step needs to support real-time stream processing to improve real-time performance and thus improve processing efficiency.
[0068] After obtaining the feature information, the epidemiological investigation information can be generated by combining the template information and the feature information. The template information can be an epidemiological investigation template. For the epidemiological investigation template, the templates corresponding to different infectious diseases can be the same or different, and the templates in different regions can also be the same or different, which can be specifically determined according to actual needs. The epidemiological investigation template can include reference feature information, such as, but not limited to, information for representing name, age, ID number, address, contact status, historical trajectory information (time and historical location), and protection information. Based on this, the feature information in the text information converted from the audio data can be automatically filled into the corresponding positions in the epidemiological investigation template to generate the epidemiological investigation information for the target object, that is, the epidemiological investigation report. In some embodiments, the filling position can be determined according to the matching relationship between the feature information and the reference feature information. After determining the filling position, the feature information can be automatically filled into the filling position in the epidemiological investigation template for automatic filling operation, so as to generate the corresponding epidemiological investigation information. Based on this, the feature information obtained by the structured service is easily combined with the epidemiological investigation template and automatically filled. The generated text can be exported in a document format with necessary styles, and thus an epidemiological investigation report with a standard and unified format can be obtained, improving the operation efficiency of generating the epidemiological investigation information.
[0069] Figure 6 The specific process diagram of conducting a telephone epidemiological investigation is schematically shown in Figure 6 As shown in
[0070] In step S610, the sending end of the interviewer establishes a dialogue request with the terminal of the target object;
[0071] In step S620, the information of the dialogue request is sent to an external device;
[0072] In step S630, the external device obtains the first input audio stream from the dialogue request and transmits it to the receiving end;
[0073] In step S640, the receiving end receives the second input audio stream input by its own microphone, obtains the first input audio stream of the dialogue request sent by the external device, combines the second input audio stream and the first input audio stream, and then generates the epidemiological investigation information by combining the mixed audio stream and the epidemiological investigation template.
[0074] In the embodiments of the present disclosure, during the process of making a dialogue request, a mixed audio stream can be generated by combining the first input audio stream of the dialogue request sent by an external device received by the receiving end and the second input audio stream input by the built-in microphone of the receiving end. Feature information is extracted from the mixed audio stream, and then, based on the feature information and the reference feature information in the template information, the extracted feature information is automatically filled into the template information to generate a flow survey report in a unified format. The technical solution in the embodiments of the present disclosure can realize automated telephone flow surveys in cooperation with mobile phones, computers, computer accessories such as audio cables and external sound cards, improve the generation efficiency of flow survey information, and reduce costs. Further, since the configuration method of the external device is simple and easy to carry, the convenience and reliability of generating flow survey information can be improved, the comprehensiveness can be improved, the application scope can be increased, which is conducive to large-scale promotion and implementation, and the usability can be improved.
[0075] In the embodiments of the present disclosure, a flow survey information processing device is further provided. As shown in Figure 7 the flow survey information processing device 700 mainly includes the following modules:
[0076] A first audio stream acquisition module 701, in response to a dialogue request between an interviewer and a target object sent by a sending end, acquires the first input audio stream in the dialogue request based on an external device;
[0077] A second audio stream acquisition module 702 is used to acquire the second input audio stream of the interviewer;
[0078] An audio stream merging module 703 is used to merge the first input audio stream and the second input audio stream to obtain a mixed audio stream;
[0079] A flow survey information generation module 704 is used to perform structured processing on the mixed audio stream to determine feature information, and determine the flow survey information of the target object in combination with the template information and the feature information.
[0080] In an exemplary embodiment of the present disclosure, the first audio stream acquisition module includes: an audio stream acquisition module, which is used to acquire audio streams of multiple audio input sources, and the audio streams include a first type identifier and a second type identifier; the first type identifier is used to determine the second input audio stream, and the second type identifier is used to determine the first input audio stream; an audio source determination module, which is used to screen the audio streams of the multiple audio input sources through the first type identifier to determine the second input audio stream from the audio streams; a first input audio stream determination module, which is used to determine the value of the second type identifier of the second input audio stream, and determine the audio stream different from the value of the second type identifier as the first input audio stream of the external device.
[0081] In an exemplary embodiment of the present disclosure, before merging the first input audio stream and the second input audio stream to obtain a mixed audio stream, the device further includes: an audio stream filtering module, configured to determine the same audio streams in the first input audio stream and / or the second input audio stream as duplicate audio streams according to a second type identifier, and perform a filtering operation on the duplicate audio streams.
[0082] In an exemplary embodiment of the present disclosure, the audio stream merging module includes: a merging control module, configured to merge the first input audio stream and the second input audio stream in chronological order to obtain the mixed audio stream.
[0083] In an exemplary embodiment of the present disclosure, before merging the first input audio stream of an external device and the second input audio stream of a receiving end to obtain a mixed audio stream, the device further includes: an audio stream adjustment module, configured to perform a sound effect adjustment operation on real-time audio parameters that meet audio conditions in the first input audio stream and the second input audio stream to adjust the first input audio stream and the second input audio stream.
[0084] In an exemplary embodiment of the present disclosure, the epidemiological investigation information generation module includes: an audio data acquisition module, configured to acquire audio data corresponding to the mixed audio stream; a template filling module, configured to convert the audio data into text information, extract feature information from the text information, and perform a filling operation on the feature information through template information to generate the epidemiological investigation information.
[0085] In an exemplary embodiment of the present disclosure, the audio data acquisition module includes: a first acquisition module, configured to, if the scenario corresponding to the dialogue request is a first type of scenario, perform a recording operation on the mixed audio stream according to recording parameters to obtain audio data that matches the speech recognition requirement; the audio parameters include one or a combination of a recording format, a sampling rate, a sampling bit depth, and a number of channels; a second acquisition module, configured to, if the scenario corresponding to the dialogue request is a second type of scenario, use the mixed audio stream as the audio data.
[0086] In addition, the specific details of each part in the above epidemiological investigation information processing device have been described in detail in the implementation manners of the epidemiological investigation information processing method part. The undisclosed detailed content can be referred to the implementation manners of the method part, and thus will not be elaborated herein.
[0087] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0088] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0089] In an embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0090] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0091] The following refers to Figure 8 to describe the electronic device 800 according to this embodiment of the present disclosure. Figure 8 The electronic device 800 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0092] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0093] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 810 can execute the steps as Figure 2 shown in
[0094] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.
[0095] The storage unit 820 may also include a program / utilities 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0096] The bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration interface, a processing unit, or a local bus using any of a variety of bus structures.
[0097] The electronic device 800 may also communicate with one or more external devices 900 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or may communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 850. Moreover, the electronic device 800 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0098] In an embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0099] A program product for implementing the above method according to an embodiment of the present disclosure may be a portable compact disc read-only memory (CD-ROM) and includes program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0100] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0101] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0102] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0103] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0104] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0105] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for processing epidemiological investigation information, characterized in that, Including: In response to a conversation request from an interviewer and a target object sent by a sending end, obtaining a first input audio stream in the conversation request based on an external device; The first input audio stream includes the audio streams of the interviewer and the target object; Obtaining a second input audio stream of the interviewer; Merging the first input audio stream and the second input audio stream to obtain a mixed audio stream; wherein, comparing the content of the first input audio stream and the second input audio stream, if the content is the same, determining the partial audio streams corresponding to the same content in the first input audio stream and the second input audio stream as duplicate audio streams, and deleting the duplicate audio streams included therein; Performing structured processing on the mixed audio stream to determine feature information, and combining template information and the feature information to determine the epidemiological investigation information of the target object.
2. The epidemiological investigation information processing method according to claim 1, wherein The obtaining the first input audio stream in the conversation request based on an external device includes: Obtaining audio streams of multiple audio input sources, the audio streams including a first type identifier and a second type identifier, wherein the first type identifier is used to determine a second input audio stream, and the second type identifier is used to determine a first input audio stream; Filtering the audio streams of the multiple audio input sources through the first type identifier to determine a second input audio stream from the audio streams; Determining the value of the second type identifier of the second input audio stream, and determining the audio stream different from the value of the second type identifier as the first input audio stream of the external device.
3. The epidemiological investigation information processing method according to claim 2, wherein Before merging the first input audio stream and the second input audio stream to obtain a mixed audio stream, the method further includes: Determining the same audio streams in the first input audio stream and / or the second input audio stream as duplicate audio streams according to the second type identifier, and performing a filtering operation on the duplicate audio streams.
4. The epidemiological investigation information processing method according to claim 1, wherein The merging the first input audio stream and the second input audio stream to obtain a mixed audio stream includes: Merging the first input audio stream and the second input audio stream in chronological order to obtain the mixed audio stream.
5. The epidemiological investigation information processing method according to claim 1, wherein Before merging the first input audio stream and the second input audio stream to obtain a mixed audio stream, the method further includes: Performing a sound effect adjustment operation on the real-time audio parameters that meet the audio conditions in the first input audio stream and the second input audio stream to adjust the first input audio stream and the second input audio stream.
6. The epidemiological investigation information processing method according to claim 1, wherein The performing structured processing on the mixed audio stream to determine feature information, and combining template information and the feature information to determine the epidemiological investigation information of the target object includes: Obtaining audio data corresponding to the mixed audio stream; Converting the audio data into text information, extracting feature information from the text information, and filling the feature information through the template information to generate the epidemiological investigation information.
7. The epidemiological investigation information processing method according to claim 6, wherein The obtaining audio data corresponding to the mixed audio stream includes: If the scenario corresponding to the dialogue request is a first - type scenario, record the mixed audio stream according to the recording parameters to obtain audio data that matches the speech recognition requirements; the audio parameters include one or a combination of recording format, sampling rate, sampling bit depth, and number of channels; If the scenario corresponding to the dialogue request is a second - type scenario, use the mixed audio stream as the audio data.
8. An epidemiological investigation information processing device, characterized in that, Comprising: A first audio stream acquisition module, configured to respond to a dialogue request between an interviewer and a target object sent by a sending end, and acquire a first input audio stream in the dialogue request based on an external device; The first input audio stream includes the audio streams of the interviewer and the target object; A second audio stream acquisition module, configured to acquire a second input audio stream of the interviewer; An audio merging module, configured to merge the first input audio stream and the second input audio stream to obtain a mixed audio stream; wherein, compare the contents of the first input audio stream and the second input audio stream, if the contents are the same, determine the partial audio streams corresponding to the same content in the first input audio stream and the second input audio stream as duplicate audio streams, and delete the duplicate audio streams included therein; An epidemiological investigation information generation module, configured to perform structured processing on the mixed audio stream to determine feature information, and determine the epidemiological investigation information of the target object in combination with template information and the feature information.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the epidemiological investigation information processing method according to any one of claims 1 - 7.
10. An electronic device, characterized in that, Comprising: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the epidemiological investigation information processing method according to any one of claims 1 - 7 by executing the executable instructions.
Citation Information
Patent Citations
Information collection and grading method and device, computer equipment and storage medium
CN111930948A
Techniques for archiving audio information
US20040121790A1