Audio processing method and device, equipment, storage medium and computer program product
By using multi-device synchronous recording, the problem of low meeting recording quality was solved, and high-quality recordings were converted into text, improving the accuracy of meeting minutes and work efficiency.
Patent Information
- Application Number
- CN202410946927.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2026-01-16
AI Technical Summary
The low recording quality during the meeting resulted in inaccurate transcripts generated from the audio recording.
By sending a recording start request to at least one second device through the first device, the first and second devices can simultaneously collect audio segments from the same sound source and generate target audio based on the audio segments at various times. The recording quality is improved by using multiple devices to record at different angles and distances.
It improves recording quality, thereby increasing the accuracy of audio-to-text transcription, helping users obtain high-quality meeting minutes and improve work efficiency.
Smart Images

Figure CN121349397A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer communications, and more particularly to an audio processing method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] Recording is a basic function of mobile phones, allowing users to capture sound, often in conversations, meetings, singing, and other scenarios. With advancements in smart technology, recording combined with audio-to-text technology has led to the ability to generate meeting minutes from audio recordings. This feature allows users to automatically generate meeting minutes without manually transcribing word for word. However, the unpredictable location of recording devices during meetings often results in low recording quality, making it difficult to accurately generate meeting minutes. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides an audio processing method, apparatus, device, storage medium, and computer program product to solve the problem of low recording quality during meetings.
[0004] According to a first aspect of the present disclosure, an audio processing method is provided, comprising:
[0005] In response to the first device initiating the audio recording function, a recording start request is sent to at least one second device; wherein the recording start request is used to instruct the second device to initiate the audio recording function;
[0006] During the audio recording process, audio segments are acquired synchronously from the same sound source by the first and second devices.
[0007] Based on the audio segments collected at various moments during the audio recording process, the first target segment corresponding to each moment is obtained;
[0008] In response to the first device stopping audio recording, target audio is generated based on the first target segment corresponding to each time point.
[0009] In some embodiments, obtaining the first target segment corresponding to each moment based on the audio segments acquired at various times during the audio recording process includes:
[0010] Based on the volume values of the audio segments acquired at the current moment, the second target segment and the third target segment are determined from the various audio segments acquired at the current moment; wherein the difference between the volume values of the second target segment and the third target segment is less than a preset threshold.
[0011] The second and third target segments are fused to obtain the first target segment corresponding to the current time.
[0012] In some embodiments, determining the second target segment and the third target segment from the various audio segments acquired at the current moment based on the volume value of the audio segments acquired at the current moment includes:
[0013] The volume values of the audio segments acquired at the current moment are compared, and the audio segment with the largest volume value among the audio segments acquired at the current moment is determined as the second target segment;
[0014] Determine the difference between the volume value of the second target segment and the volume values of each audio segment other than the second target segment;
[0015] Audio segments with a difference less than a preset threshold are identified as the third target segment.
[0016] In some embodiments, the method further includes:
[0017] The target audio is converted into text to obtain the audio processing result, which is then output in text format.
[0018] In some embodiments, the method further includes:
[0019] In response to the first device entering the audio recording interface, a joint recording request is sent to at least one second device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0020] Obtain confirmation information from at least one second device regarding the joint recording request; wherein the confirmation information is used to indicate whether to accept or reject the joint recording request;
[0021] The step of sending a recording start request to at least one second device in response to the first device starting the audio recording function includes:
[0022] In response to the first device initiating the audio recording function, a recording start request is sent to the second device, which indicates acceptance of the joint recording request.
[0023] In some embodiments, the first device includes at least two audio acquisition modules, and the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0024] According to a second aspect of the present disclosure, an audio processing method is also provided, comprising:
[0025] The second device receives a recording start request from the first device;
[0026] Based on the request to start recording, the audio recording function of the second device is started, and audio segments are synchronously acquired from the same sound source while the first device is recording audio.
[0027] The audio segments captured at various moments during the audio recording process are used by the first device to obtain the first target segment corresponding to each moment, and when the first device stops the audio recording function, the target audio is generated based on the first target segment corresponding to each moment.
[0028] In some embodiments, the method further includes:
[0029] Receive a joint recording request from the first device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0030] Send a confirmation message to the first device in response to the joint recording request; wherein the confirmation message is used to indicate whether to accept the joint recording request or to refuse to accept the joint recording request;
[0031] The second device receives a recording start request from the first device, including:
[0032] If the confirmation message indicates that the second device accepts the joint recording request, the first device accepts the request to start recording.
[0033] The recording start request is sent by the first device when the audio recording function is started.
[0034] In some embodiments, the second device includes at least two audio acquisition modules, wherein the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0035] According to a third aspect of the present disclosure, an audio processing apparatus is provided, comprising:
[0036] The sending module is configured to send a recording start request to at least one second device in response to the first device starting the audio recording function; wherein the recording start request is used to instruct the second device to start the audio recording function;
[0037] The first acquisition module is configured to acquire audio segments synchronously collected from the same sound source by the first device and the second device during the audio recording process.
[0038] The second acquisition module is configured to obtain the first target segment corresponding to each moment based on the audio segments collected at each moment during the audio recording process.
[0039] The generation module is configured to generate target audio based on the first target segment corresponding to each time point in response to the first device stopping audio recording.
[0040] In some embodiments, the apparatus further includes:
[0041] Based on the volume values of the audio segments acquired at the current moment, the second target segment and the third target segment are determined from the various audio segments acquired at the current moment; wherein the difference between the volume values of the second target segment and the third target segment is less than a preset threshold.
[0042] The second and third target segments are fused to obtain the first target segment corresponding to the current time.
[0043] In some embodiments, the apparatus further includes:
[0044] The volume values of the audio segments acquired at the current moment are compared, and the audio segment with the largest volume value among the audio segments acquired at the current moment is determined as the second target segment;
[0045] Determine the difference between the volume value of the second target segment and the volume values of each audio segment other than the second target segment;
[0046] Audio segments with a difference less than a preset threshold are identified as the third target segment.
[0047] In some embodiments, the apparatus further includes:
[0048] The target audio is converted into text to obtain the audio processing result, which is then output in text format.
[0049] In some embodiments, the apparatus further includes:
[0050] In response to the first device entering the audio recording interface, a joint recording request is sent to at least one second device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0051] Obtain confirmation information from at least one second device regarding the joint recording request; wherein the confirmation information is used to indicate whether to accept or reject the joint recording request;
[0052] The step of sending a recording start request to at least one second device in response to the first device starting the audio recording function includes:
[0053] In response to the first device initiating the audio recording function, a recording start request is sent to the second device, which indicates acceptance of the joint recording request.
[0054] In some embodiments, the first device protected on the device side includes at least two audio acquisition modules, and the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0055] According to a fourth aspect of the present disclosure, an audio processing apparatus is also provided, comprising:
[0056] The receiving module is configured so that the second device receives a recording start request from the first device;
[0057] The acquisition module is configured to initiate the audio recording function of the second device based on the recording start request, and to synchronously acquire audio segments from the same sound source with the first device during the audio recording process of the first device.
[0058] The audio segments captured at various moments during the audio recording process are used by the first device to obtain the first target segment corresponding to each moment, and when the first device stops the audio recording function, the target audio is generated based on the first target segment corresponding to each moment.
[0059] In some embodiments, the apparatus further includes:
[0060] Receive a joint recording request from the first device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0061] Send a confirmation message to the first device in response to the joint recording request; wherein the confirmation message is used to indicate whether to accept the joint recording request or to refuse to accept the joint recording request;
[0062] The second device receives a recording start request from the first device, including:
[0063] If the confirmation message indicates that the second device accepts the joint recording request, the first device accepts the request to start recording.
[0064] The recording start request is sent by the first device when the audio recording function is started.
[0065] In some embodiments, the second device protected on the device side includes at least two audio acquisition modules, wherein the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0066] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0067] processor;
[0068] Memory used to store computer programs or instructions;
[0069] The processor executes the computer program or instructions to implement the steps of the method described in any one of the first and second aspects above.
[0070] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, the storage medium storing a computer program or instructions that, when executed by a processor, implement the steps of the method described in any one of the first and second aspects.
[0071] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program or instructions, which, when executed by a processor, implement the steps of the method described in any one of the first and second aspects.
[0072] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0073] In the audio processing method proposed in this embodiment, in response to the first device starting the audio recording function, a recording start request is sent to at least one second device; during the audio recording process, audio segments synchronously collected from the same sound source by the first device and the second device are acquired; based on the audio segments collected at various times during the audio recording process, a first target segment corresponding to each time moment is obtained; in response to the first device stopping the audio recording function, a target audio is generated based on the first target segment corresponding to each time moment.
[0074] In other words, when the first device starts audio recording, it sends a recording start request to at least one second device. During recording, the first and second devices simultaneously capture audio segments from the same sound source. Based on the audio segments captured at different times, a first target segment for each time moment is determined. When the first device stops audio recording, the target audio is finally generated using the first target segments corresponding to each time moment. By combining multiple devices to record simultaneously from different angles and distances, recording quality can be improved, thereby increasing the accuracy of audio-to-text transcription, helping users obtain high-quality meeting minutes, and improving work efficiency.
[0075] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0076] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0077] Figure 1 This is a flowchart illustrating an audio processing method according to an exemplary embodiment.
[0078] Figure 2 This is a flowchart illustrating another audio processing method according to an exemplary embodiment.
[0079] Figure 3 This is a timing diagram illustrating an audio processing method according to an exemplary embodiment.
[0080] Figure 4 This is a spatial schematic diagram of a first device and a second device according to an exemplary embodiment.
[0081] Figure 5 This is a block diagram illustrating an audio processing apparatus according to an exemplary embodiment.
[0082] Figure 6 This is a block diagram illustrating another audio processing apparatus according to an exemplary embodiment.
[0083] Figure 7 This is a structural block diagram of an apparatus 700 according to an exemplary embodiment.
[0084] Figure 8 This is a block diagram illustrating an audio processing apparatus 800 according to an exemplary embodiment. Detailed Implementation
[0085] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0086] Figure 1 This is a flowchart illustrating an audio processing method according to an exemplary embodiment. Figure 1 As shown, the method mainly includes the following steps:
[0087] In step 101, in response to the first device starting the audio recording function, a recording start request is sent to at least one second device; wherein the recording start request is used to instruct the second device to start the audio recording function;
[0088] In step 102, during the audio recording process, audio segments are acquired synchronously from the same sound source by the first device and the second device.
[0089] In step 103, the first target segment corresponding to each moment is obtained based on the audio segments collected at various times during the audio recording process;
[0090] In step 104, in response to the first device stopping the audio recording function, target audio is generated based on the first target segment corresponding to each time moment.
[0091] It should be noted that the audio processing method proposed in this disclosure can be applied to electronic devices. Here, electronic devices can include terminal devices, such as mobile terminals or fixed terminals. Mobile terminals can include devices such as mobile phones, tablets, laptops, and wearable electronic devices. Fixed terminals can include desktop computers, smart TVs, and in-vehicle systems.
[0092] It should be noted that the first device is the master device responsible for recording the meeting, playing a leading role throughout the recording process. The second device, as a slave device, works in conjunction with the master device, recording from different distances and angles to provide more comprehensive and multi-dimensional sound capture. Furthermore, the master device also has control functions for the slave device, allowing remote control to start or stop recording, thus ensuring flexibility and synchronization in the recording process. This master-slave collaborative recording method improves the recording quality.
[0093] It should be noted that the second device may include one or more. The first device may send a request via broadcast to devices that meet the auxiliary recording conditions, requesting these devices to join the joint recording. Devices that meet the auxiliary recording conditions may decide whether to join the joint recording as second devices based on actual circumstances and needs. Devices that meet the auxiliary recording conditions must have recording functionality and may include at least one of the following: devices located at a distance less than a preset distance from the first device, devices in the same wireless network environment as the first device, user-preset devices, or devices with preset recording quality or functionality.
[0094] Understandably, to improve recording quality, both the first and second devices can be equipped with at least two audio acquisition modules. These modules can be microphones with noise reduction capabilities, allowing for more effective sound capture and reception while minimizing ambient noise interference. Furthermore, to facilitate subsequent processing of the acquired audio clips, both devices need data transmission capabilities to ensure efficient transfer of audio clips between them.
[0095] It should be noted that data transmission between the first and second devices can be achieved through at least one of the following methods: local area network (LAN) connection, Bluetooth connection, Wi-Fi Direct connection, or near-field communication (NFC) connection. For example, audio and recording requests can be transmitted via LAN or Bluetooth, and fast data exchange can be performed via Wi-Fi Direct or NFC. The first and second devices can establish a Bluetooth or LAN connection via broadcast.
[0096] It should be noted that when a trigger operation targeting the recording application is detected, the audio recording interface will be entered. When a startup operation targeting the audio recording interface is detected, the audio recording function can be started.
[0097] Understandably, when the second device receives a request to start recording, it can verify the request. If it verifies that the request was indeed sent by the first device, the second device will start the audio recording function based on the request.
[0098] It should be noted that after receiving a request to start recording, the second device will verify the request: first, it will confirm whether the request format is correct; then, it will verify the identity of the request sender through technologies such as digital signatures and check the sender's permissions. In addition, it will check the uniqueness of the request and verify whether the request has been tampered with. Only after all verifications are successful will the second device start the audio recording function.
[0099] It should be noted that the first and second devices simultaneously collect audio clips from the same sound source, and the second device will send the audio clips collected at each moment to the first device in real time via Bluetooth or local area network connection.
[0100] It should be noted that, in order to synchronously acquire sound from the same sound source, the first and second devices first synchronize their times to ensure that their clocks are accurately aligned. Then, the first device sends a synchronization start signal to the second device, triggering both to begin audio acquisition simultaneously. Furthermore, because the first and second devices are physically close to each other, use the same recording parameters, and are located in the same space, they are able to acquire sound from the same source.
[0101] It should be noted that the sound source can be the voices of the participants in the meeting, or the sound of music, videos, or other media played on electronic devices during the meeting.
[0102] Understandably, the sound acquisition by the first and second devices is synchronized. For example, if the first device acquires an audio segment at 14:34:51, the second device will also acquire that same segment simultaneously. Once the second device acquires the audio segment, it immediately sends it to the first device via the communication connection. Due to the time required for data transmission, the first device might acquire the audio segment from the second device at 14:34:52. The first device will immediately perform a fusion process upon receiving the audio segment from the second device; it will not cache the audio segment acquired by the second device to avoid consuming excessive storage space.
[0103] It is understandable that audio clips can be captured every second (s), every 5 seconds, or every 5 minutes (min). The specific capture interval can be determined based on the actual situation and needs.
[0104] It should be noted that, based on the audio segments captured at various moments during the audio recording process, the first target segment corresponding to each moment can be obtained. This first target segment is derived by processing the audio segments captured at each moment. For example, the audio segment with the highest waveform among the audio segments captured at each moment can be used as the first target segment for that moment. Alternatively, the audio segments captured at each moment can be further subdivided, and the audio segments with the highest sound quality among the subdivided segments can be merged, with the merged audio segment serving as the first target segment for that moment. If multiple audio segments are captured at the current moment, the audio segment with the highest volume among these multiple audio segments can be determined as the first target segment for that moment, thus obtaining the first target segment for each moment.
[0105] Understandably, generating target audio by using the first target segment acquired and processed by the first and second devices at different times can effectively avoid sound quality problems caused by environmental noise, signal interference, or limitations of the devices themselves. This not only reduces noise and distortion in the audio but also ensures the clarity and purity of the audio signal, thereby improving the overall recording quality. Furthermore, high-quality recording can enhance the accuracy of audio-to-text transcription and reduce the error rate in the conversion process.
[0106] In the audio processing method proposed in this embodiment, in response to the first device starting the audio recording function, a recording start request is sent to at least one second device; during the audio recording process, audio segments synchronously collected from the same sound source by the first device and the second device are acquired; based on the audio segments collected at various times during the audio recording process, a first target segment corresponding to each time moment is obtained; in response to the first device stopping the audio recording function, a target audio is generated based on the first target segment corresponding to each time moment.
[0107] In other words, when the first device starts audio recording, it sends a recording start request to at least one second device. During recording, the first and second devices simultaneously capture audio segments from the same sound source. Based on the audio segments captured at different times, a first target segment for each time moment is determined. When the first device stops audio recording, the target audio is finally generated using the first target segments corresponding to each time moment. By combining multiple devices to record simultaneously from different angles and distances, recording quality can be improved, thereby increasing the accuracy of audio-to-text transcription, helping users obtain high-quality meeting minutes, and improving work efficiency.
[0108] In some embodiments, obtaining the first target segment corresponding to each moment based on the audio segments acquired at various times during the audio recording process includes:
[0109] Based on the volume values of the audio segments acquired at the current moment, the second target segment and the third target segment are determined from the various audio segments acquired at the current moment; wherein the difference between the volume values of the second target segment and the third target segment is less than a preset threshold.
[0110] The second and third target segments are fused to obtain the first target segment corresponding to the current time.
[0111] It should be noted that the acquired audio segments can be parsed through the audio parsing interface to obtain the volume value of the audio segments.
[0112] Understandably, when the volume difference between two audio clips is less than a preset threshold, the two audio clips can be considered to be of similar recording quality. Since the two audio clips are collected from different angles and distances, they can be fused using a sound enhancement algorithm, which can improve the overall recording quality and make the recording clearer and fuller.
[0113] For example, after obtaining each audio segment collected at the current moment, the difference between the volume values of any two audio segments in each audio segment can be determined, thereby obtaining multiple differences. Then, the two audio segments with differences less than a preset threshold are determined as the second target segment and the third target segment. In other embodiments, when there are multiple differences less than the preset threshold, the candidate audio segment with the larger volume value among the two audio segments corresponding to each difference can be determined, and the volume values of the multiple candidate audio segments corresponding to the multiple differences can be compared. The candidate audio segment with the largest volume value is determined as the second target segment, and the audio segment with the difference from the candidate audio segment is determined as the third target segment.
[0114] In other embodiments, when there are multiple differences less than a preset threshold, the differences can be compared, and the two audio segments corresponding to the minimum difference can be determined as the second target segment and the third target segment, respectively.
[0115] In other embodiments, when there are multiple differences less than a preset threshold, multiple sets of two audio segments with differences less than the preset threshold can be identified as the second target segment and the third target segment, respectively. This allows multiple sets of second target segments and third target segments to be identified and then fused together.
[0116] In other embodiments, after obtaining each audio segment collected at the current moment, the difference between the volume values of any two audio segments in each audio segment can be determined to obtain multiple differences. Then, the differences are sorted, and the two audio segments corresponding to the minimum value of each difference are determined as the second target segment and the third target segment, respectively.
[0117] In other embodiments, the audio segment with the highest volume among the various audio segments collected at the current moment can be determined as the second target segment, and the audio segment whose volume difference with the second target segment is less than a preset threshold can be determined as the third target segment.
[0118] In other embodiments, any audio segment among the various audio segments acquired at the current moment can be determined as the second target segment, and audio segments whose volume difference with the second target segment is less than a preset threshold can be determined as the third target segment. After obtaining the second and third target segments, methods such as adaptive filters, spectral subtraction, and wavelet domain denoising algorithms can be used to fuse audio segments acquired at different distances and angles, thereby obtaining the first target segment after noise reduction or enhancement.
[0119] In this embodiment, a second target segment and a third target segment can be selected from the audio segments captured at the current moment based on their volume values, wherein the volume difference between the second and third target segments is within a preset threshold. Subsequently, the second and third target segments are fused to generate the first target segment corresponding to the current moment. By fusing audio segments with similar volumes, a clearer and higher-quality audio segment can be obtained.
[0120] In some embodiments, determining the second target segment and the third target segment from the various audio segments acquired at the current moment based on the volume value of the audio segments acquired at the current moment includes:
[0121] The volume values of the audio segments acquired at the current moment are compared, and the audio segment with the largest volume value among the audio segments acquired at the current moment is determined as the second target segment;
[0122] Determine the difference between the volume value of the second target segment and the volume values of each audio segment other than the second target segment;
[0123] Audio segments with a difference less than a preset threshold are identified as the third target segment.
[0124] It's understandable that the volume of an audio clip is directly proportional to its quality; the higher the volume, the higher the audio quality. Therefore, the first step is to select the audio clip with the highest volume, which is the second target clip. Then, based on the second target clip, one or more third target clips similar to it are found. Finally, by fusing the second and third target clips, a higher quality audio clip can be obtained.
[0125] Understandably, since all audio segments with a difference less than a preset threshold will be identified as the third target segment, there can be multiple third target segments.
[0126] For example, a preset threshold of 10 dB is set. Five audio segments are currently collected: A, B, C, D, and E. The volume values of these five segments are 30 dB, 34 dB, 60 dB, 54 dB, and 57 dB, respectively. By comparing the volume values of these five segments, the segment with the highest volume, C, is identified as the second target segment. Next, the differences between the volume value of audio segment C and those of audio segments A, B, D, and E are calculated. The results show differences of 30 dB, 26 dB, 6 dB, and 3 dB, respectively. By comparing these differences with the preset threshold, it is found that the volume differences between audio segments D and E and the volume value of audio segment C are less than the preset threshold. Therefore, audio segments D and E can be identified as the third target segments.
[0127] In other embodiments, after determining the difference between the volume value of the second target segment and the volume values of each audio segment other than the second target segment, the differences can be directly sorted, and the audio segment corresponding to the minimum value among the differences can be determined as the third target segment.
[0128] In this embodiment, by comparing the volume values of audio segments acquired at the current moment, the segment with the highest volume is identified as the second target segment. Next, the volume difference between the second target segment and all other audio segments is calculated, and segments with a difference less than a preset threshold are identified as the third target segment. This process can identify audio segments with prominent volume as well as other audio segments with similar volume, which is helpful for feature extraction, noise reduction, or audio enhancement in subsequent audio processing, thereby improving the recording quality.
[0129] In some embodiments, the method further includes:
[0130] The target audio is converted into text to obtain the audio processing result, which is then output in text format.
[0131] It should be noted that text conversion processing can be achieved through text recognition algorithms. These algorithms include Optical Character Recognition (OCR), Natural Language Processing (NLP), Automatic Speech Recognition (ASR), and deep learning techniques such as Recurrent Neural Networks (RNNs) or Convolutional Neural Networks (CNNs). By using text recognition algorithms, the speech content in target audio can be accurately converted into text, helping users obtain accurate meeting transcripts.
[0132] In this embodiment of the disclosure, by performing accurate text conversion processing on the target audio, the audio content can be obtained efficiently and output clearly in text format, which can help users obtain high-quality meeting records and improve work efficiency.
[0133] In some embodiments, the method further includes:
[0134] In response to the first device entering the audio recording interface, a joint recording request is sent to at least one second device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0135] Obtain confirmation information from at least one second device regarding the joint recording request; wherein the confirmation information is used to indicate whether to accept or reject the joint recording request;
[0136] The step of sending a recording start request to at least one second device in response to the first device starting the audio recording function includes:
[0137] In response to the first device initiating the audio recording function, a recording start request is sent to the second device, which indicates acceptance of the joint recording request.
[0138] It should be noted that when the second device receives a joint recording request, it will display the request content on its interface and provide the user with operation controls. The user can accept the joint recording request by clicking the confirmation control or reject it by clicking the rejection control. Both actions will send corresponding confirmation information to the first device. By obtaining the user's confirmation information, the user's intentions can be accurately understood, and corresponding operational adjustments can be made to ensure the smooth progress of the joint recording process or to terminate it in a timely manner.
[0139] Understandably, if the confirmation message from the second device indicates that it refuses to accept the joint recording request, it means that the second device is processing other tasks and cannot perform audio recording operations. In this case, when the first device starts the audio recording function, it does not need to send a start recording request to the second device. If the confirmation message from the second device indicates that it accepts the joint recording request, it means that the second device is in an idle state and can perform audio recording operations. In this case, when the first device starts the audio recording function, it can send a start recording request to the second device.
[0140] In this embodiment of the disclosure, when the first device enters the audio recording interface, the first device sends a joint recording request to at least one second device, instructing the second device to participate in the audio recording of the first device. After receiving confirmation information from the second device regarding the joint recording request, the first device determines whether the second device accepts the joint recording request based on the confirmation information. When the first device starts the audio recording function, the first device sends a start recording request to the second device that has confirmed acceptance of the joint recording request, to ensure that multiple parties can simultaneously record audio.
[0141] In some embodiments, the first device includes at least two audio acquisition modules, and the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0142] For example, taking a first device equipped with two audio acquisition modules as an example, the first audio acquisition module can be set on one side of the device, with the audio acquisition surface facing a major sound source, such as the front; the second audio acquisition module can be set on the other side of the device, with the audio acquisition surface facing another important sound source, such as the rear or side.
[0143] For example, taking the first device equipped with three audio acquisition modules as an example, the first audio acquisition module can be set at the front of the device, mainly responsible for acquiring the sound in front; the second audio acquisition module can be set at the rear or side of the device, responsible for acquiring the sound from different directions to increase the sense of space; the third audio acquisition module can be set at the top or bottom of the device as needed, used to capture the sound above or below, further improving the stereo effect and improving the sound quality.
[0144] In this embodiment of the disclosure, the first device is equipped with at least two audio acquisition modules, each with its audio acquisition surface facing a different direction, thereby enabling it to capture sound from all directions and providing richer, more multi-angle sound source capture capabilities for audio recording.
[0145] Figure 2 This is a flowchart illustrating another audio processing method according to an exemplary embodiment. Figure 2 As shown, the method mainly includes the following steps:
[0146] In step 201, the second device receives a request to start recording from the first device;
[0147] In step 202, based on the request to start recording, the audio recording function of the second device is started, and audio segments are synchronously acquired from the same sound source while the first device is recording audio.
[0148] The audio segments captured at various moments during the audio recording process are used by the first device to obtain the first target segment corresponding to each moment, and when the first device stops the audio recording function, the target audio is generated based on the first target segment corresponding to each moment.
[0149] It should be noted that the audio processing method proposed in this disclosure can be applied to electronic devices. Here, electronic devices can include terminal devices, such as mobile terminals or fixed terminals. Mobile terminals can include devices such as mobile phones, tablets, laptops, and wearable electronic devices. Fixed terminals can include desktop computers, smart TVs, and in-vehicle systems.
[0150] It should be noted that the first device is the master device responsible for recording the meeting, playing a leading role throughout the recording process. The second device, as a slave device, works in conjunction with the master device, recording from different distances and angles to provide more comprehensive and multi-dimensional sound capture. Furthermore, the master device also has control functions for the slave device, allowing remote control to start or stop recording, thus ensuring flexibility and synchronization in the recording process. This master-slave collaborative recording method improves the recording quality.
[0151] It should be noted that the second device may include one or more. The first device may send a request via broadcast to devices that meet the auxiliary recording conditions, requesting these devices to join the joint recording. Devices that meet the auxiliary recording conditions may decide whether to join the joint recording as second devices based on actual circumstances and needs. Devices that meet the auxiliary recording conditions must have recording functionality and may include at least one of the following: devices located at a distance less than a preset distance from the first device, devices in the same wireless network environment as the first device, user-preset devices, or devices with preset recording quality or functionality.
[0152] Understandably, to improve recording quality, both the first and second devices can be equipped with at least two audio acquisition modules. These modules can be microphones with noise reduction capabilities, allowing for more effective sound capture and reception while minimizing ambient noise interference. Furthermore, to facilitate subsequent processing of the acquired audio clips, both devices need data transmission capabilities to ensure efficient transfer of audio clips between them.
[0153] It should be noted that data transmission between the first and second devices can be achieved through at least one of the following methods: local area network (LAN) connection, Bluetooth connection, Wi-Fi Direct connection, or near-field communication (NFC) connection. For example, audio and recording requests can be transmitted via LAN or Bluetooth, and fast data exchange can be performed via Wi-Fi Direct or NFC. The first and second devices can establish a Bluetooth or LAN connection via broadcast.
[0154] Understandably, when the second device receives a request to start recording, it can verify the request. If it verifies that the request was indeed sent by the first device, the second device will start the audio recording function based on the request.
[0155] It should be noted that after receiving a request to start recording, the second device will verify the request: first, it will confirm whether the request format is correct; then, it will verify the identity of the request sender through technologies such as digital signatures and check the sender's permissions. In addition, it will check the uniqueness of the request and verify whether the request has been tampered with. Only after all verifications are successful will the second device start the audio recording function.
[0156] It should be noted that the first and second devices simultaneously collect audio clips from the same sound source, and the second device will send the audio clips collected at each moment to the first device in real time via Bluetooth or local area network connection.
[0157] It should be noted that, in order to synchronously acquire sound from the same sound source, the first and second devices first synchronize their times to ensure that their clocks are accurately aligned. Then, the first device sends a synchronization start signal to the second device, triggering both to begin audio acquisition simultaneously. Furthermore, because the first and second devices are physically close to each other, use the same recording parameters, and are located in the same space, they are able to acquire sound from the same source.
[0158] It should be noted that the sound source can be the voices of the participants in the meeting, or the sound of music, videos, or other media played on electronic devices during the meeting.
[0159] Understandably, the sound acquisition by the first and second devices is synchronized. For example, if the first device acquires an audio segment at 14:34:51, the second device will also acquire that same segment simultaneously. Once the second device acquires the audio segment, it immediately sends it to the first device via the communication connection. Due to the time required for data transmission, the first device might acquire the audio segment from the second device at 14:34:52. The first device will immediately perform a fusion process upon receiving the audio segment from the second device; it will not cache the audio segment acquired by the second device to avoid consuming excessive storage space.
[0160] It is understandable that audio clips can be captured every second (s), every 5 seconds, or every 5 minutes (min). The specific capture interval can be determined based on the actual situation and needs.
[0161] It should be noted that, based on the audio segments captured at various moments during the audio recording process, the first target segment corresponding to each moment can be obtained. This first target segment is derived by processing the audio segments captured at each moment. For example, the audio segment with the highest waveform among the audio segments captured at each moment can be used as the first target segment for that moment; alternatively, the audio segments captured at each moment can be further subdivided, and the audio segments with the highest sound quality among the subdivided segments can be merged, with the merged audio segment serving as the first target segment for that moment.
[0162] Understandably, generating target audio by using the first target segment acquired and processed by the first and second devices at different times can effectively avoid sound quality problems caused by environmental noise, signal interference, or limitations of the devices themselves. This not only reduces noise and distortion in the audio but also ensures the clarity and purity of the audio signal, thereby improving the overall recording quality. Furthermore, high-quality recording can enhance the accuracy of audio-to-text transcription and reduce the error rate in the conversion process.
[0163] In this embodiment, upon receiving a recording start request from the first device, the second device immediately activates its own audio recording function, synchronously acquiring audio segments from the same sound source as the first device. During audio recording, the audio segments acquired at various times are used by the first device to obtain the corresponding first target segment. When the first device stops recording, it uses the acquired first target segment to generate the final target audio. This multi-device synchronous recording method not only improves the flexibility and quality of audio acquisition but also provides richer material for target audio generation through multi-angle synchronous recording, thereby enhancing the efficiency and quality of audio production.
[0164] In some embodiments, the method further includes:
[0165] Receive a joint recording request from the first device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0166] Send a confirmation message to the first device in response to the joint recording request; wherein the confirmation message is used to indicate whether to accept the joint recording request or to refuse to accept the joint recording request;
[0167] The second device receives a recording start request from the first device, including:
[0168] If the confirmation message indicates that the second device accepts the joint recording request, the first device accepts the request to start recording.
[0169] The recording start request is sent by the first device when the audio recording function is started.
[0170] In this embodiment of the disclosure, after receiving a joint recording request from the first device, the second device sends a confirmation message to indicate whether it accepts the joint recording request. Once accepted, the second device will synchronously receive a recording start request when the first device starts audio recording, thereby realizing joint recording by the first and second devices. This not only expands the scope of audio recording collaboration but also improves the flexibility and quality of recording through inter-device cooperation, bringing more possibilities to audio production and enhancing the reliability and diversity of the recording process.
[0171] In some embodiments, the second device includes at least two audio acquisition modules, wherein the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0172] For example, taking the second device equipped with two audio acquisition modules as an example, the first audio acquisition module can be set on one side of the device, with the audio acquisition surface facing a major sound source, such as the front; the second audio acquisition module can be set on the other side of the device, with the audio acquisition surface facing another important sound source, such as the rear or side.
[0173] For example, taking the second device equipped with three audio acquisition modules, the first audio acquisition module is set at the front of the device and is mainly responsible for acquiring the sound from the front; the second audio acquisition module is set at the rear or side of the device and is responsible for acquiring the sound from different directions to increase the sense of space; the third audio acquisition module can be set at the top or bottom of the device as needed to capture the sound from above or below, further improving the stereo effect and sound quality.
[0174] In this embodiment of the disclosure, the second device is equipped with at least two audio acquisition modules, each with its audio acquisition surface facing a different direction, thereby enabling it to capture sound from all directions and providing richer, more multi-angle sound source capture capabilities for audio recording.
[0175] Figure 3 This is a timing diagram illustrating an audio processing method according to an exemplary embodiment, such as... Figure 3 As shown, the main steps include:
[0176] In step 301, in response to starting the audio recording function, a recording start request is sent to at least one second device;
[0177] In step 302, a request to start recording is received from the first device;
[0178] In step 303, the audio recording function is started based on the request to start recording;
[0179] In step 304, an audio segment synchronously acquired from the same sound source is sent to the first device;
[0180] In step 305, the first target segment corresponding to each moment is obtained based on the audio segments collected at various times during the audio recording process;
[0181] In step 306, in response to the stop audio recording function, target audio is generated based on the first target segment corresponding to each time point.
[0182] Figure 4 This is a spatial schematic diagram of a first device and a second device according to an exemplary embodiment, such as... Figure 4 As shown, a first device, a second device A, and a second device B synchronously acquire audio segments from the same sound source. The first device acquires audio segments directly in front of the sound source, while devices A and B acquire segments at 45-degree angles above and below the sound source, respectively. After acquiring the audio segments, devices A and B immediately transmit them to the first device for processing. The first device then fuses the acquired audio segments to obtain the target audio. By acquiring audio segments from multiple angles and distances, and then generating the target audio based on these segments, the quality and accuracy of the audio can be effectively improved.
[0183] Figure 5 This is a block diagram illustrating an audio processing apparatus according to an exemplary embodiment. Figure 5 As shown, the device mainly includes:
[0184] The sending module 501 is configured to send a recording start request to at least one second device in response to the first device starting the audio recording function; wherein the recording start request is used to instruct the second device to start the audio recording function;
[0185] The first acquisition module 502 is configured to acquire audio segments synchronously collected from the same sound source by the first device and the second device during the audio recording process.
[0186] The second acquisition module 503 is configured to obtain the first target segment corresponding to each moment based on the audio segments collected at each moment during the audio recording process.
[0187] The generation module 504 is configured to generate target audio based on the first target segment corresponding to each time point in response to the first device stopping audio recording.
[0188] In some embodiments, the device 500 further includes:
[0189] Based on the volume values of the audio segments acquired at the current moment, the second target segment and the third target segment are determined from the various audio segments acquired at the current moment; wherein the difference between the volume values of the second target segment and the third target segment is less than a preset threshold.
[0190] The second and third target segments are fused to obtain the first target segment corresponding to the current time.
[0191] In some embodiments, the device 500 further includes:
[0192] The volume values of the audio segments acquired at the current moment are compared, and the audio segment with the largest volume value among the audio segments acquired at the current moment is determined as the second target segment;
[0193] Determine the difference between the volume value of the second target segment and the volume values of each audio segment other than the second target segment;
[0194] Audio segments with a difference less than a preset threshold are identified as the third target segment.
[0195] In some embodiments, the device 500 further includes:
[0196] The target audio is converted into text to obtain the audio processing result, which is then output in text format.
[0197] In some embodiments, the device 500 further includes:
[0198] In response to the first device entering the audio recording interface, a joint recording request is sent to at least one second device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0199] Obtain confirmation information from at least one second device regarding the joint recording request; wherein the confirmation information is used to indicate whether to accept or reject the joint recording request;
[0200] The step of sending a recording start request to at least one second device in response to the first device starting the audio recording function includes:
[0201] In response to the first device initiating the audio recording function, a recording start request is sent to the second device, which indicates acceptance of the joint recording request.
[0202] In some embodiments, the first device protected by the device side 500 includes at least two audio acquisition modules, and the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0203] Figure 6 This is a block diagram of another audio processing apparatus according to an exemplary embodiment. Figure 6 As shown, the device mainly includes:
[0204] The receiving module 601 is configured to receive a recording start request from the first device;
[0205] The acquisition module 602 is configured to start the audio recording function of the second device based on the start recording request, and to synchronously acquire audio segments from the same sound source with the first device during the audio recording process of the first device;
[0206] The audio segments captured at various moments during the audio recording process are used by the first device to obtain the first target segment corresponding to each moment, and when the first device stops the audio recording function, the target audio is generated based on the first target segment corresponding to each moment.
[0207] In some embodiments, the device 600 further includes:
[0208] Receive a joint recording request from the first device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device;
[0209] Send a confirmation message to the first device in response to the joint recording request; wherein the confirmation message is used to indicate whether to accept the joint recording request or to refuse to accept the joint recording request;
[0210] The second device receives a recording start request from the first device, including:
[0211] If the confirmation message indicates that the second device accepts the joint recording request, the first device accepts the request to start recording.
[0212] The recording start request is sent by the first device when the audio recording function is started.
[0213] In some embodiments, the second device protected by the device side 600 includes at least two audio acquisition modules, and the audio acquisition surfaces of each audio acquisition module are oriented differently.
[0214] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0215] Figure 7 This is a structural block diagram illustrating an apparatus 700 according to an exemplary embodiment. For example, apparatus 700 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0216] Reference Figure 7 The device 700 may include one or more of the following components: processing component 702, memory 704, power supply component 706, multimedia component 708, audio component 710, input / output (I / O) interface 712, sensor component 714, and communication component 716.
[0217] Processing component 702 typically controls the overall operation of device 700, such as operations associated with at least one of display, telephone call, data communication, camera operation, and recording operation. Processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.
[0218] Memory 704 is configured to store various types of data to support operation on device 700. Examples of such data include at least one of the following: instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, and videos. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0219] Power supply component 706 provides power to various components of device 700. Power supply component 706 may include at least one of the following: a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 700.
[0220] Multimedia component 708 includes a screen that provides an output interface between device 700 and the user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen may be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0221] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.
[0222] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, and buttons. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0223] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or one of its components, the presence or absence of user contact with device 700, orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include an optical sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include, but is not limited to, at least one of the following: an accelerometer, a gyroscope, a magnetometer, a pressure sensor, and a temperature sensor.
[0224] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards, such as Wi-Fi, 4G, 5G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.
[0225] In an exemplary embodiment, the device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components.
[0226] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including executable instructions or a computer program, which can be executed by the processor 720 of the device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0227] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform any of the audio processing methods described above in the embodiments of this disclosure. For example, the method includes:
[0228] In response to the first device initiating the audio recording function, a recording start request is sent to at least one second device; wherein the recording start request is used to instruct the second device to initiate the audio recording function;
[0229] During the audio recording process, audio segments are acquired synchronously from the same sound source by the first and second devices.
[0230] Based on the audio segments collected at various moments during the audio recording process, the first target segment corresponding to each moment is obtained;
[0231] In response to the first device stopping audio recording, target audio is generated based on the first target segment corresponding to each time point.
[0232] This disclosure provides a computer program product comprising a computer program or executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or executable instructions from the computer-readable storage medium and executes the computer program or executable instructions, causing the computer device to perform any of the audio processing methods described above in this disclosure.
[0233] Figure 8 This is a block diagram illustrating an apparatus 800 for audio processing according to an exemplary embodiment. For example, apparatus 800 may be provided as a server. (Refer to...) Figure 8 The device 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by memory 832 for storing instructions, such as application programs, that can be executed by the processing component 822. The application programs stored in memory 832 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 822 is configured to execute instructions to perform any of the aforementioned audio processing methods. For example, the method includes:
[0234] In response to the first device initiating the audio recording function, a recording start request is sent to at least one second device; wherein the recording start request is used to instruct the second device to initiate the audio recording function;
[0235] During the audio recording process, audio segments are acquired synchronously from the same sound source by the first and second devices.
[0236] Based on the audio segments collected at various moments during the audio recording process, the first target segment corresponding to each moment is obtained;
[0237] In response to the first device stopping audio recording, target audio is generated based on the first target segment corresponding to each time point.
[0238] Device 800 may also include a power supply component 826 configured to perform power management of device 800, a wired or wireless network interface 850 configured to connect device 800 to a network, and an input / output (I / O) interface 858. Device 800 can operate an operating system stored in memory 832, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0239] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0240] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An audio processing method, characterized by, The method comprises: sending a start recording request to at least one second device in response to the first device starting an audio recording function; wherein the start recording request is used to instruct the second device to start the audio recording function; during the audio recording process, obtaining audio clips collected synchronously by the first device and the second device from the same sound source; based on the audio clips collected at each moment during the audio recording process, obtaining a first target clip corresponding to each moment; in response to the first device stopping the audio recording function, generating a target audio based on the first target clip corresponding to each moment.
2. The method of claim 1, wherein, The method comprises: based on the volume value of the audio clip collected at the current moment, determining a second target clip and a third target clip from the audio clips collected at the current moment; wherein the difference between the volume value of the second target clip and the volume value of the third target clip is less than a preset threshold value; performing fusion processing on the second target clip and the third target clip to obtain the first target clip corresponding to the current moment.
3. The method of claim 2, wherein, The method comprises: comparing the volume values of the audio clips collected at the current moment, and determining the audio clip with the largest volume value as the second target clip; determining the difference between the volume value of the second target clip and the volume values of the audio clips other than the second target clip; determining the audio clip with a difference less than the preset threshold value as the third target clip.
4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: performing text conversion processing on the target audio to obtain an audio processing result, and outputting the audio processing result in text format.
5. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: sending a joint recording request to at least one second device in response to the first device entering an audio recording interface; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device; obtaining confirmation information of the at least one second device for the joint recording request; wherein the confirmation information is used to indicate acceptance of the joint recording request, or rejection of the joint recording request; The method comprises: in response to the first device starting the audio recording function, sending the start recording request to the second device whose confirmation information indicates acceptance of the joint recording request.
6. The method according to any one of claims 1 to 3, characterized in that, The first device comprises: at least two audio collection modules, and the orientations of the audio collection surfaces of each audio collection module are different.
7. An audio processing method, characterized by, The method comprises: The second device receives a start recording request from the first device; based on the start recording request, starting the audio recording function of the second device, and collecting audio clips synchronously with the first device from the same sound source during the audio recording process of the first device; The audio segments collected at each moment during the audio recording process are used by the first device to obtain first target segments corresponding to each moment, and based on the first target segments corresponding to each moment, the target audio is generated in the case that the first device stops the audio recording function.
8. The method of claim 7, wherein, The method further comprises: receiving a joint recording request from the first device; wherein the joint recording request is used to instruct the second device to participate in the audio recording of the first device; sending confirmation information for the joint recording request to the first device; wherein the confirmation information is used to indicate acceptance of the joint recording request, or rejection of the joint recording request; The second device receives a start recording request from the first device, comprising: in the case that the confirmation information indicates that the second device accepts the joint recording request, accepting the start recording request from the first device; wherein the start recording request is sent by the first device in the case that the audio recording function is started.
9. The method according to claim 7 or 8, characterized in that, The second device comprises at least two audio acquisition modules, and the orientations of the audio acquisition surfaces of each audio acquisition module are different.
10. An audio processing apparatus, characterized by comprising: Comprising: a sending module configured to send a start recording request to at least one second device in response to the first device starting the audio recording function; wherein the start recording request is used to instruct the second device to start the audio recording function; a first obtaining module configured to obtain audio segments synchronously collected from the same sound source by the first device and the second device during the audio recording process; a second obtaining module configured to obtain first target segments corresponding to each moment based on the audio segments collected at each moment during the audio recording process; a generating module configured to generate target audio based on the first target segments corresponding to each moment in response to the first device stopping the audio recording function.
11. An audio processing apparatus, characterized by comprising: Comprising: a receiving module configured to receive a start recording request from the first device by the second device; an acquisition module configured to start the audio recording function of the second device based on the start recording request, and to synchronously collect audio segments from the same sound source with the first device during the audio recording process by the first device; wherein the audio segments collected at each moment during the audio recording process are used by the first device to obtain first target segments corresponding to each moment, and based on the first target segments corresponding to each moment, the target audio is generated in the case that the first device stops the audio recording function.
12. An electronic device, comprising: Comprising: a processor; a memory for storing computer programs or instructions; wherein the processor executes the computer programs or instructions to implement the steps of the method of any one of claims 1 to 9.
13. A non-transitory computer-readable storage medium storing a computer program or instructions, wherein, When the computer programs or instructions in the storage medium are executed by the processor, the steps of the method of any one of claims 1 to 9 are implemented.
14. A computer program product comprising computer programs or instructions, characterized in that, The computer programs or instructions are executed by the processor to implement the steps of the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for mobile terminal synergistic shooting
CN104426588A
Sound recording method and device
CN105261385A
Information processing method and system, first equipment and second equipment
CN110334240A
Bluetooth voice audio acquisition method and system
CN111370012A
Recording system and recording method
CN111986715A