Audio processing method and device, recording terminal, electronic equipment and storage medium

By performing face tracking and voiceprint feature matching on the video stream of the business hall, the problem of audio segmentation not being automated in existing technologies has been solved. This enables audio stream segmentation based on individuals, improving the accuracy and efficiency of finding audio files while protecting customer privacy.

CN116486788BActive Publication Date: 2026-01-02IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310466334.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2026-01-02
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

Existing technology cannot automatically segment audio based on customer business transactions, making it impossible to quickly and accurately locate specific recording files.

Method used

By performing face tracking on the video stream of the target area, and using timing or changes in personnel identification to automatically segment the audio stream, and combining voiceprint feature matching to merge similar audio segments, audio stream segmentation is achieved on a per-person basis.

Benefits of technology

It enables automated audio stream segmentation, improving the accuracy and efficiency of audio file retrieval and protecting customer privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486788B_ABST
    Figure CN116486788B_ABST
Patent Text Reader

Abstract

The application provides an audio processing method, device, recording terminal, electronic equipment and storage medium, by acquiring a video stream corresponding to a target area, performing face tracking based on the video stream, and starting timing in the case that tracking of a single first person identifier or all first person identifiers fails, until the timing duration reaches a preset duration, or the tracking result changes to a second person identifier, the audio stream corresponding to the target area is split based on the timing duration, or the first person identifier and the second person identifier, an automatic audio stream splitting scheme based on personnel is realized, the defects that the existing scheme can only perform timing splitting or manual operation splitting on the audio, so that the splitting method cannot distinguish the customers corresponding to the audio or is not high in effectiveness are overcome, and conditions are provided for quickly and accurately finding a specified audio file according to actual needs subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, and in particular to an audio processing method and device, a recording terminal, electronic equipment and a storage medium. BACKGROUND

[0002] In the process of business handling, the customer and the clerk are usually recorded in order to facilitate the application of subsequent related information, such as quality inspection and evaluation of the service of the clerk, and finding a specified recording file. In order to quickly and accurately find the specified recording file, the audio of the business hall is usually cut and processed.

[0003] At present, when the audio of the business hall is cut and processed, the audio of the customer and the clerk is usually collected by an audio collection device, and the collected audio is cut and processed by a timing cutting scheme and then saved, or the clerk manually operates the start and stop of the audio collection device. However, the existing scheme cannot automatically cut the audio according to the business handling customer, so it is impossible to quickly and accurately find a specified recording file of a customer as needed. SUMMARY

[0004] The present application provides an audio processing method, device, recording terminal, electronic equipment and storage medium, which solves the problem that the prior art cannot cut the audio according to the business handling customer, and realizes an automatic audio stream cutting scheme based on personnel.

[0005] The present application provides an audio processing method, comprising:

[0006] Obtaining a video stream corresponding to a target area, and performing face tracking based on the video stream;

[0007] In a case where the tracking result of the face tracking includes a single first personnel identifier, tracking the single first personnel identifier in the video stream until the tracking of the single first personnel identifier fails, and starting a timer;

[0008] In a case where the tracking result of the face tracking includes a plurality of first personnel identifiers, tracking the plurality of first personnel identifiers in the video stream until the tracking of all first personnel identifiers fails, and starting a timer;

[0009] Continuously timing until the timing duration reaches a preset duration, or the tracking result changes to a second personnel identifier;

[0010] Cutting an audio stream corresponding to the target area based on the timing duration, or the first personnel identifier and the second personnel identifier.

[0011] The audio processing method provided in the present application comprises the following steps:

[0012] In the case that the tracking result is changed to the second personnel identifier, the audio stream is divided based on the first personnel identifier and the second personnel identifier.

[0013] In the case that the tracking result is changed to the second personnel identifier, the audio stream is divided based on the first personnel identifier and the second personnel identifier.

[0014] The audio processing method provided in the present application comprises the following steps:

[0015] In the case that the tracking result is changed to the second personnel identifier, the audio stream is divided based on the first personnel identifier and the second personnel identifier.

[0016] In the case that the tracking result is changed to the second personnel identifier, the audio stream is divided based on the first personnel identifier and the second personnel identifier.

[0017] The audio processing method provided in the present application comprises the following steps:

[0018] The plurality of first audio segments obtained by dividing are acquired.

[0019] The first audio segments with matching voiceprint features are merged to obtain a plurality of second audio segments.

[0020] The audio processing method provided in the present application comprises the following steps:

[0021] The first audio segments with matching voiceprint features are merged to obtain a plurality of second audio segments.

[0022] The audio processing method provided in the present application comprises the following steps:

[0023] The video stream is subjected to personnel detection.

[0024] In a case where the detection result of the person detection indicates that there is a person, face tracking is performed on the video stream, and the audio stream is collected.

[0025] According to the audio processing method provided by the application, the preset time length is determined based on the interval time length between adjacent sample audio segments obtained by cutting the sample audio stream.

[0026] The application further provides an audio processing device, comprising:

[0027] A tracking unit is configured to acquire a video stream corresponding to a target area, and perform face tracking based on the video stream.

[0028] A first starting unit is configured to, in a case where the tracking result of the face tracking includes a single first person identifier, track the single first person identifier in the video stream until the tracking of the single first person identifier fails, and start timing.

[0029] A second starting unit is configured to, in a case where the tracking result of the face tracking includes a plurality of first person identifiers, track the plurality of first person identifiers in the video stream until the tracking of the plurality of first person identifiers all fails, and start timing.

[0030] A timing unit is configured to continuously time until the timing time length reaches a preset time length, or the tracking result changes to a second person identifier.

[0031] A cutting unit is configured to cut an audio stream corresponding to the target area based on the timing time length, or the first person identifier and the second person identifier.

[0032] The application further provides a recording terminal, comprising:

[0033] A microphone is configured to collect an audio stream corresponding to a target area.

[0034] A camera is configured to collect a video stream corresponding to the target area.

[0035] A processor connected with the microphone and the camera respectively, configured to acquire the audio stream and the video stream, and perform face tracking based on the video stream; in a case that the tracking result of the face tracking includes a single first person identifier, track the single first person identifier in the video stream until the tracking of the single first person identifier fails, and start timing, or in a case that the tracking result of the face tracking includes a plurality of first person identifiers, track the plurality of first person identifiers in the video stream until the tracking of all the first person identifiers fails, and start timing; continue timing until the timing duration reaches a preset duration, or the tracking result changes to a second person identifier; and cut the audio stream based on the timing duration, or the first person identifier and the second person identifier.

[0036] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the audio processing method according to any one of the preceding embodiments when executing the program.

[0037] The application further provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the audio processing method according to any one of the preceding embodiments.

[0038] The application provides an audio processing method, device, recording terminal, electronic device and storage medium, which performs face tracking on a video stream of a target area, and starts timing in a case that the tracking of a single first person identifier or all the first person identifiers fails, thereby cutting an audio stream corresponding to the target area based on a timing duration, or based on a first person identifier and a second person identifier obtained by changing the tracking result, to realize an automatic audio stream cutting scheme based on personnel, and overcome the defects in the prior art that the audio can only be cut in a timing manner or manually, so that the cutting method cannot distinguish the customers corresponding to the audio or is not effective, thereby providing conditions for quickly and accurately finding a specified audio file according to actual needs. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without creative effort.

[0040] Figure 1 is one of the flowcharts of the audio processing method provided by the application;

[0041] Figure 2Fig. 2 is a flowchart of the audio processing method provided by the present application;

[0042] Figure 3 Fig. 3 is a flowchart of step 110 of the audio processing method provided by the present application;

[0043] Figure 4 Fig. 4 is a structural diagram of the audio processing device provided by the present application;

[0044] Figure 5 Fig. 5 is a structural diagram of the recording terminal provided by the present application;

[0045] Figure 6 Fig. 6 is a scene diagram of the customer handling business in the business hall provided by the present application;

[0046] Figure 7 Fig. 7 is a structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0047] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0048] At present, when the audio in the business hall is divided, the audio of the customer and the clerk is generally collected by the audio collection device, and the collected audio is divided by a timing division scheme, or the start and stop of the audio collection device is manually operated by the clerk. The timing division scheme cannot automatically divide the audio according to the business handling customer, so that the corresponding customer of the audio cannot be distinguished, and thus the audio file of the specified customer cannot be accurately found. Although the start and stop of the audio collection device can be manually operated by the clerk to divide the audio according to the business handling customer, it needs to be manually operated by the clerk, and cannot realize automatic division of the audio, which reduces the efficiency of audio division, and also increases the working intensity of the clerk. To this end, the present application provides an audio processing method to overcome the above problems.

[0049] Figure 1 Fig. 1 is a flowchart of the audio processing method provided by the present application, which includes the following steps: Figure 1

[0050] Step 110, acquiring a video stream corresponding to a target area, and performing face tracking based on the video stream;

[0051] ​It should be noted that the target area is a pre-defined area which needs to be collected and processed by audio segmentation. When the audio processing method provided by the embodiment of the present application is applied to the segmentation processing of the business communication audio of the customers and the clerks collected in the business hall, the target area refers to an area corresponding to the target customer according to the position of the target customer when conducting the business, and the target customer refers to the customer conducting the business at the counter or the window of the business hall.

[0052] Specifically, before acquiring the video stream corresponding to the target area, the target area needs to be determined first, and the target area can be set according to the effective sound pickup area in the embodiment of the present application. When the initialization setting of the audio segmentation in the business hall is performed, the effective sound pickup area and the suppression area can be determined according to the sound production positions of the clerks and the customers initialized at the counter or the window, and the target area is set according to the effective sound pickup area after the effective sound pickup area is determined.

[0053] After the target area is determined, the video stream corresponding to the target area can be acquired in real time by a video acquisition device, such as a camera. It should be understood that in the scenario of the audio segmentation in the business hall, the face or the body area of the target customer conducting the business can be collected in the video stream corresponding to the target area.

[0054] Face tracking refers to the process of capturing the position and the size of the face and other information in the subsequent frames on the premise that the face is detected, including the face recognition and the face tracking technology. In the embodiment of the present application, after the video stream corresponding to the target area is acquired, the specific face target needs to be tracked based on the video stream, and the specific face target refers to the target customer entering the target area. Obviously, to track the face target in the image, the face needs to be recognized first, and the face recognition is to find out the face from the static picture or the video sequence and output the number, the position and the size of the face and other effective information by using the computer. Then the face is tracked on the premise that the face is detected, and the position and the size of the face and other information are captured in the subsequent frames. In the embodiment of the present application, the face tracking technology is used to track the target customer, and the identity information corresponding to the recognized face does not need to be acquired, which can effectively protect the privacy of the customer.

[0055] Step 120, in the case that the tracking result of the face tracking includes a single first personnel identifier, tracking the single first personnel identifier in the video stream until the tracking of the single first personnel identifier is invalid, and starting the timing;

[0056] Step 130, in the case that the tracking result of the face tracking includes a plurality of first personnel identifiers, tracking the plurality of first personnel identifiers in the video stream until the tracking of all the first personnel identifiers is invalid, and starting the timing;

[0057] It should be noted that when face tracking is performed based on the video stream, the face in the image is first recognized and the number, position and size of the face and other effective information are output. In general, the number of faces output by face recognition can be zero, one or more. On the premise that the face is detected, i.e., the face recognition result is one or more, the face is continuously captured and tracked. When the number of faces output by face recognition is zero, it indicates that no face is detected in the video stream. When the number of faces output by face recognition is one, the face tracking continues to track and capture the single face target in the subsequent frames. When the number of faces output by face recognition is more than one, the face tracking needs to continue to track and capture the multiple face targets in the subsequent frames.

[0058] Specifically, the first personnel identifier in steps 120 and 130 represents the face target detected and recognized when face tracking is performed based on the video stream. In the case where the tracking result of face tracking includes a single first personnel identifier, it indicates that the number of detected faces is one, indicating that there is a single target customer in the target area at this time. The single target customer, i.e., the first personnel identifier, is continuously tracked until the tracking of the single first personnel identifier fails, where the tracking failure indicates that no face is detected and recognized in the image, indicating that the single target customer in the target area has left at this time.

[0059] In the case where the tracking result of face tracking includes multiple first personnel identifiers, it indicates that the number of detected faces is more than one, indicating that there are multiple target customers in the target area at this time. The multiple target customers, i.e., the first personnel identifiers, are continuously tracked until the tracking of all first personnel identifiers fails, indicating that the multiple target customers in the target area have all left at this time.

[0060] Considering the situation that a customer leaves and then returns in the actual scene, the timing can be started in the case where the tracking of the single first personnel identifier fails or the tracking of the multiple first personnel identifiers all fails, i.e., in the case where the single target customer leaves the target area or the multiple target customers all leave the target area, thereby counting the duration of the target customer leaving the target area. It can be understood that in the case where the tracking result includes multiple first personnel identifiers, the multiple first personnel identifiers are tracked in the video stream until the tracking of all first personnel identifiers fails, indicating that the last target customer in the multiple target customers leaves the target area, at which time the timing is started.

[0061] Step 140, the timing is continuously performed until the duration of the timing reaches a preset duration, or the tracking result changes from empty to a second personnel identifier;

[0062] Specifically, the second personnel identifier represents a face target detected and recognized based on face tracking of the video stream, and the tracking result of the face tracking is the second personnel identifier, which covers the case where the target customer is one or more. During the timing, the timing duration is constantly updated. When the timing duration reaches the preset duration, it indicates that the duration for which the target customer leaves the target area exceeds the preset duration, at which time it can be considered that the target customer has completed the business handling and left the counter or window of the business hall, and in this case, the timing can be exited. It should be understood that the preset duration is a pre-set duration threshold, which refers to the general interval duration for adjacent different target customers to handle business, and the specific value thereof can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0063] When the tracking result of the face tracking is the second personnel identifier, it indicates that one or more target customers are detected in the target area at this time to handle business. This one or more target customers can be the target customer who temporarily left during the business handling process and then returned to continue handling, or can be a new target customer, that is, the first personnel identifier and the second personnel identifier can be consistent, can be different, or can have partial overlap. As long as the tracking result changes to the second personnel identifier, it indicates that the business handling process is re-entered, and at this time the timing can be exited.

[0064] In the embodiments of the present application, the timing can be exited when the timing duration reaches the preset duration or when the tracking result changes to the second personnel identifier. The timing duration reaching the preset duration is to exit the timing under the condition of meeting the timing, and the tracking result changing from null to the second personnel identifier is to detect the second personnel identifier when the timing has not ended, so the timing can also be exited.

[0065] Step 150, based on the timing duration, or the first personnel identifier and the second personnel identifier, the audio stream corresponding to the target area is segmented.

[0066] It should be noted that before the audio segmentation, the business hall audio to be segmented needs to be obtained, which can usually be collected by an audio collection device. The audio collection device of the embodiments of the present application can record and suppress sound at a specified angle, so when initializing, the effective pickup area and the suppression area can be determined according to the pronunciation positions of the customers and the clerks initialized at the counter or window, and after the effective pickup area is determined, the audio stream in the effective pickup area can be obtained through the audio collection device. In addition, after the effective pickup area is determined, the target area can be set correspondingly according to the effective pickup area, so the video stream and the audio stream obtained are both for the same target area.

[0067] Specifically, after obtaining the audio stream corresponding to the target area, the audio stream needs to be cut, and in the embodiment of the application, the audio stream corresponding to the target area is cut based on the timing duration or the first personnel identifier and the second personnel identifier. The audio file obtained by cutting can facilitate subsequent quick and accurate search for the audio file of the specified customer.

[0068] In the audio cutting process, whether the cutting is performed based on the timing duration or the first personnel identifier and the second personnel identifier is determined based on the condition met by the timing in step 140.

[0069] For example, in the case where the timing duration reaches the preset duration in step 140, the audio stream is cut based on the timing duration, that is, the time when the timing is exited is counted back by the timing duration to obtain the time when the target customer corresponding to the first personnel identifier leaves the target area, that is, the time when the tracking of the single first personnel identifier or all first personnel identifiers fails, as the end time of the audio file corresponding to the first personnel identifier, and the audio stream is cut. For another example, in the case where the tracking result of the face tracking changes to the second personnel identifier in step 140, the first personnel identifier and the second personnel identifier are compared to determine whether the target customer goes away and returns or a new target customer goes to conduct business, and then it is determined whether the time when the tracking of the single first personnel identifier or all first personnel identifiers fails is taken as the end time of the audio file corresponding to the first personnel identifier or whether the audio stream is cut after the tracking of all second personnel identifiers fails.

[0070] The audio processing method provided by the embodiment of the application realizes the automatic audio stream cutting scheme in units of personnel by tracking the face of the target area and starting the timing in the case where the tracking of the single first personnel identifier or all first personnel identifiers fails, thereby cutting the audio stream corresponding to the target area based on the timing duration or based on the first personnel identifier and the second personnel identifier obtained by changing the tracking result. The method overcomes the defect in the prior art that the audio can only be cut in a timing manner or manually, so that the cutting method cannot distinguish the customers corresponding to the audio or is not effective, and provides a condition for subsequent quick and accurate search for the specified audio file according to actual needs.

[0071] Based on the above embodiment, step 150 includes step 151 and step 152, and step 151 and step 152 are executed according to actual conditions:

[0072] Step 151, in the case where the timing duration reaches the preset duration, the audio stream is cut based on the first time, and the first time is the time when the tracking of the single first personnel identifier or all first personnel identifiers fails.

[0073] Specifically, the timing duration refers to the duration of the target customer leaving the target area, and when the timing duration reaches the preset duration, it indicates that the duration of the target customer leaving the target area exceeds the preset duration, at which time it is usually considered that the target customer has completed the business handling and left the counter or window of the business hall, and in this case, the audio stream can be segmented to obtain the audio file corresponding to the target customer.

[0074] In the embodiment of the application, the audio stream is segmented based on the first time after the target customer completes the business handling and leaves the target area. The first time refers to the time when the tracking of the single first person identifier or all first person identifiers fails, that is, the specific time when the single target customer or all target customers leave the target area after completing the business handling. Segmenting the audio stream based on the first time not only can accurately obtain the audio file corresponding to the target customer, but also can make the segmented audio segment not contain a silent segment, so that the audio segmentation is more effective.

[0075] Step 152, in the case where the tracking result changes to the second person identifier, segmenting the audio stream based on the first person identifier and the second person identifier.

[0076] Specifically, when the tracking result of the face tracking is a single or multiple first person identifiers, it indicates that a single or multiple target customers are detected in the target area to handle the business at this time, and the single or multiple target customers can be marked as first target customers. When the tracking of the single first person identifier or all first person identifiers fails, it indicates that the first target customer has left the target area. When the tracking result of the face tracking changes to the second person identifier, it indicates that one or more target customers are detected in the target area to handle the business at this time, and the target customer is marked as a second target customer. Since the second target customer and the first target customer can be the same or different customers, the audio stream can be segmented based on the same or different first person identifier and second person identifier, realizing an automated audio stream segmentation scheme based on personnel, and providing conditions for subsequent quick and accurate search for a specified audio file.

[0077] Based on the above embodiment, the second target customer and the first target customer can be different customers or the same customer. When the first target customer temporarily leaves the target area (the timing duration does not exceed the preset duration) during the business handling and then returns to continue the business, the second target customer and the first target customer can be considered as the same customer handling the business. For this, the embodiment of the application proposes a method for segmenting the audio stream corresponding to the target area in the case where the first target customer and the second target customer are the same customer or different customers. Step 152 specifically includes step 1521 and step 1522, and step 1521 and step 1522 are executed according to actual conditions:

[0078] Step 1521, in the case where there is a single first personnel identification and the single first personnel identification is consistent with the second personnel identification, or in the case where there are multiple first personnel identifications and there is at least one first personnel identification consistent with the second personnel identification, return to face tracking;

[0079] Step 1522, in the case where there is a single first personnel identification and the single first personnel identification is different from the second personnel identification, or in the case where there are multiple first personnel identifications and the multiple first personnel identifications are all different from the second personnel identification, based on the first time, split the audio stream.

[0080] Specifically, in the case where the tracking result is changed to the second personnel identification, it is necessary to compare the first personnel identification and the second personnel identification, so as to determine whether there is a same personnel identification in the first personnel identification and the second personnel identification.

[0081] In the case where the first personnel identification is single, i.e., the number of the first target customers is one, the first personnel identification and the second personnel identification are consistent, at this time, whether the number of the second target customers is one or multiple, it means that there is an intersection between the first target customers and the second target customers, i.e., the first target customers return, at this time, no audio stream splitting is performed, and the face tracking is continued until the tracking of the single first personnel identification is re-failed, and then the timing is returned.

[0082] In the case where the first personnel identification is multiple, i.e., the number of the first target customers is multiple, there is at least one first personnel identification consistent with the second personnel identification in the first personnel identification, at this time, whether the number of the second target customers is one or multiple, it means that there is an intersection between the first target customers and the second target customers, i.e., at least one first target customer returns, at this time, no audio stream splitting is performed, and the face tracking is continued until the tracking of the multiple first personnel identifications is re-failed, and then the timing is returned.

[0083] In the case where the single first personnel identification is different from the second personnel identification or the multiple first personnel identifications are all different from the second personnel identification, it means that there is no same customer between the first target customers and the second target customers, i.e., whether the number of the first target customers and the second target customers is one or multiple, there is no intersection between the first target customers and the second target customers, i.e., there is no customer returning, at this time, the timing can be jumped out, and the face tracking is continued until the tracking of all the second personnel identifications is failed, and then the timing is restarted, and at this time, the audio stream is split based on the first time, which refers to the time when the tracking of the single first personnel identification or all the first personnel identifications is failed, i.e., the time when the first target customers leave the target area.

[0084] Further, the first personnel identification is single, indicating that one target customer in the target area is detected to handle the business; the first personnel identification is multiple, indicating that multiple target customers in the target area are detected to handle the business. When the target customer in the target area is detected to handle the business, the single target customer can be marked as A1, the tracking result of the face tracking is obtained, the first personnel identification is tracked, and the face tracking of A1 is continued. When A1 leaves the target area, the timing is started, and the time length of A1 leaving the target area is counted. When A1 returns to the target area within the preset time length, that is, the time length of A1 leaving does not exceed the preset time length, the face target marked in the tracking result of the face tracking of the second personnel identification is still A1, that is, the first personnel identification and the second personnel identification are consistent. In this case, it is generally considered that A1 temporarily leaves and then returns to continue to handle the business, and it is considered as the same customer to continue to handle the business. Therefore, in the case that the first personnel identification and the second personnel identification are consistent, the timing is exited, and the face tracking of A1 is continued to ensure that the audio file corresponding to the target customer A1 obtained by finally cutting is complete.

[0085] When A1 leaves the target area, the timing is started, and the time length of A1 leaving the target area is counted. In the process of timing, although the timing time length has not reached the preset time length, the face tracking detects a new target customer in the target area again, and the tracking result of the second personnel identification is B1. Since A1 and B1 are different face targets marked, the first personnel identification and the second personnel identification are different. In this case, it can be considered that different target customers enter the target area to handle the business. Therefore, in the case that the first personnel identification and the second personnel identification are different, the audio stream in the target area can be cut based on the first time, that is, the time when the tracking of the single first personnel identification fails, that is, the time when the customer A1 leaves the target area. The audio stream is cut based on the time, and the audio file corresponding to the customer A1 can be obtained without containing a mute segment, so that the audio cutting is more effective.

[0086] When it is detected that there are multiple target customers in the target region, the multiple target customers can be marked as A1, …, An (where n≥2) respectively, a tracking result first personnel identifier of face tracking is obtained, face tracking is continuously performed on the multiple target customers, when one target customer, such as A1, leaves the target region, timing is not started, only when all the multiple target customers leave the target region, timing is started, and timing is started at the moment when the last target customer in the multiple target customers leaves, when at least one target customer in the multiple target customers returns to the target region before the timing duration reaches the preset duration, the timing is exited, and face tracking is continued on the target customer. Because there is at least one first personnel identifier consistent with the second personnel identifier in the first personnel identifier, it can be considered that the same target customer is still handling the business, and the audio stream is not divided at this time.

[0087] When all the multiple target customers leave the target region, although the timing duration does not reach the preset duration, face tracking detects that a new target customer enters the target region again during the timing, and a tracking result second personnel identifier obtained can be B1 or B1, …, Bn (where n≥2), A1, …, An and B1, …, Bn are different faces marked, and thus the second personnel identifier is different from the first personnel identifier. In this case, it can be considered that different target customers enter the target region to handle the business, and thus in the case where the first personnel identifier and the second personnel identifier are different, the audio stream in the target region is divided based on a first moment, the first moment refers to a moment when tracking of all the first personnel identifiers fails, that is, a moment when the last customer in the customers A1, …, An leaves the target region, the audio stream is divided based on the moment, an audio file corresponding to the target customer can be obtained, and a mute segment is not included, so that the audio division is more effective.

[0088] Based on any of the above embodiments, Figure 2 A flowchart of an audio processing method provided in an embodiment of the present application is shown in Figure 2. In the embodiment of the present application, considering the case where the timing duration reaches the preset duration based on the first moment to divide the audio stream provided in the above embodiment, the division method does not consider the case where a customer leaves the target region beyond the preset duration and then returns to continue handling. For example, a customer suddenly needs to leave temporarily during the process of handling the business, although the customer returns to continue handling the business subsequently, the duration of leaving the target region exceeds the preset duration, resulting in that the audio file corresponding to the customer is divided into two segments, so that the specified audio file cannot be accurately found subsequently. Although face tracking is performed on the customer in the target region, in order to protect the privacy of the customer, the face information cannot be stored, so that the audio file of the customer cannot be combined based on the face information subsequently.

[0089] To solve the problem, the embodiment of the present application provides a method, by comparing the voiceprint features of the audio segments, the audio segments with the same voiceprint features are merged, thereby obtaining the audio files in the unit of customers. As shown in the figure, the method further comprises the following steps after step 150: Figure 2

[0090] Step 160, obtaining the plurality of first audio segments obtained by the segmentation;

[0091] Step 170, merging the first audio segments with the same voiceprint features in the plurality of first audio segments, thereby obtaining the plurality of second audio segments.

[0092] Specifically, the plurality of first audio segments refers to the plurality of audio segments obtained by segmenting the audio stream corresponding to the target area according to the audio processing method provided in the above embodiment. Considering that some customers may temporarily leave and then return to continue the business, and since the duration of their leaving the target area exceeds the preset duration, the audio file corresponding to the customer is segmented. In order to merge the audio file of this type of customer, all audio segments corresponding to the target area obtained by segmentation on the same day are usually obtained. After obtaining the plurality of first audio segments obtained by segmentation, the voiceprint features of each first audio segment can be extracted, and then the voiceprint features of each first audio segment are matched according to the similarity between the voiceprint features, thereby merging the first audio segments with the same voiceprint features, i.e. the first audio segments corresponding to the same customer, and obtaining the plurality of second audio segments in the unit of customers.

[0093] Further, the similarity between the voiceprint features can reflect the similarities and differences between the customers corresponding to different first audio segments. In the case that the similarity between the voiceprint features of two first audio segments reaches a preset similarity threshold, it can be determined that the two first audio segments are the audio of the same customer, and therefore the two first audio segments can be merged, and the merged audio segment is recorded as a second audio segment. If the similarity between the voiceprint features of two first audio segments does not reach the preset similarity threshold, it is determined that the two first audio segments belong to the audio segments of different customers, and are not merged, but are directly taken as second audio segments. It should be understood that the preset similarity threshold is preset, and is used to determine the similarity value of the voiceprint features indicating that the two first audio segments correspond to the same person, which can be set according to actual needs, and the embodiment of the present application does not make specific limitation.

[0094] ​In the embodiment of the present application, the voiceprint features of the plurality of first audio segments obtained by segmentation are compared, the first audio segments with matched voiceprint features are merged to obtain a plurality of second audio segments, the effect of merging the audio segments corresponding to the same customer is realized, the audio file in the unit of customer can be obtained, and the problem that the audio corresponding to the customer is divided into two segments due to the duration of the customer leaving the target area exceeding the preset duration is effectively solved. Moreover, the merging manner based on the matching of the voiceprint features does not need to apply the face information, and the privacy of the customer can be effectively protected.

[0095] Based on the above embodiment, step 170 comprises: performing voiceprint feature matching on adjacent first audio segments in the plurality of first audio segments, and merging the matched first audio segments.

[0096] Specifically, in the business hall audio segmentation scenario, when the customer returns, generally, no new customer will be inserted to handle business in the middle, therefore, when performing voiceprint feature comparison on the plurality of first audio segments, the embodiment of the present application performs voiceprint feature comparison on all adjacent first audio segments in the plurality of first audio segments. That is, the voiceprint feature matching in the embodiment of the present application is only performed on the first audio segments adjacent in time sequence.

[0097] In this way, not only the audio segments corresponding to the same customer can be merged, but also the calculation amount caused by matching all first audio segments can be reduced, and the efficiency of audio merging can be improved.

[0098] Based on any of the above embodiments, Figure 3 The flowchart of step 110 in the audio processing method provided by the embodiment of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, in step 110, face tracking is performed based on the video stream, and specifically comprises:

[0099] Step 111, detecting personnel in the video stream;

[0100] Step 112, in the case that the detection result of the personnel detection indicates that there is personnel, performing face tracking on the video stream and collecting the audio stream.

[0101] Specifically, since the face tracking is performed on the premise of detecting the face, after obtaining the video stream corresponding to the target area, personnel detection needs to be performed on the video stream to determine whether there is a specific personnel target in the target area.

[0102] It can be understood that the personnel detection herein can be realized by face detection or by infrared detection, etc. Taking face detection as an example, the personnel detection on the video stream is essentially a process of face image collection and detection. When a customer enters the target area, i.e. the customer is within the shooting range of the collection device, the collection device will automatically search and shoot the face image of the customer, and automatically mark the position and size of the face in the image.

[0103] When the detection result of the personnel detection indicates the presence of personnel, it means that a customer has entered the target area, and the customer is the target customer. After the face of the target customer is detected, the face information of the target customer is continuously captured in the subsequent frames, and the audio collection device is started to collect the audio of the customer and the salesperson, so as to obtain the audio stream corresponding to the target area.

[0104] In the embodiment of the present application, the condition that the detection result of the personnel detection indicates the presence of personnel is taken as a trigger condition to perform face tracking on the video stream and collect the audio stream. Not only the audio stream corresponding to the target area can be collected, but also the audio collection device can be prevented from being turned on for a long time, and the collected audio stream can be prevented from containing a large number of invalid silent segments, so that the audio collection is more efficient and accurate.

[0105] Based on any of the above embodiments, the preset time length is determined based on the interval time length between adjacent sample audio segments obtained by cutting the sample audio stream.

[0106] It should be noted that the sample refers to a part of individuals observed or investigated. The sample audio stream in the embodiment of the present application can be selected from the audio streams of several business halls, and the sample audio segment can be an audio segment obtained by manually operating the start and stop of the audio collection device according to different business handling customers by the salesperson.

[0107] Specifically, the preset time length is a pre-set time threshold, which can be set to 8s, 10s, 15s, etc. It can be determined based on the interval time length between sample audio segments obtained by cutting the sample audio stream. When the sample audio segment is obtained by manually cutting by the salesperson, the interval time length between adjacent sample audio segments in the business hall can be statistically analyzed to obtain the sample mean and sample variance of the interval time length, and a suitable preset time length can be determined based on the sample mean and sample variance of the interval time length.

[0108] Based on any of the above embodiments, the present embodiment provides an audio processing method for realizing business hall audio cutting processing in units of customers. The method comprises:

[0109] Detect whether a customer enters a target area, if so, start the audio acquisition device to record the customer and the clerk, and record the business handling start time T1. The customer in the target area is collected face image, based on the collected image to determine how many customers are currently handling business at the same time, and mark the current customer, such as detecting a single customer handling business, the single customer is marked as A1, such as detecting two customers handling business at the same time, the two customers can be marked as A1, B1.

[0110] In the case of detecting a single customer handling business, the single customer A1 is continuously tracked, and if A1 is detected to leave, the timing starts, and if A1 is detected to return within a preset time period, the face tracking of A1 continues; if A1 is not detected to return after the preset time period, the timing ends, the detection of A1 stops, the business handling start time T1 and the business handling end time T2 of the customer A1 are recorded, and the audio file in the time period is saved together.

[0111] In the case of detecting multiple customers handling business, the multiple customers are continuously tracked, and in the embodiment of the application, two customers handling business at the same time are taken as an example for introduction. The two customers are marked as A1 and B1 respectively, if one of them is detected to leave, the timing does not start; when A1 and B1 are detected to leave, the timing starts, if A1 or B1 is detected to return within a preset time period, the face tracking of the returned customer continues; if at least one of A1 and B1 is not detected to return after the preset time period, the timing ends, the detection of A1 and B1 stops, the business handling start time T1 and the business handling end time T2 corresponding to the customers A1 and B1 are recorded, and the audio file in the time period is saved together.

[0112] After the business time of the day ends, all the audio files L1~Ln of the customers of the day are obtained, the voiceprint detection processing is performed on the above audio files, the voiceprint features of each adjacent two audio files are compared, for example, L1 and L2 are compared, L2 and L3 are compared……, the audio files with consistent voiceprint features are merged into an audio file of a customer, and the merged audio file is saved.

[0113] Based on any of the above embodiments, Figure 4 As shown in the structure schematic diagram of the audio processing device provided by the embodiment of the application, the device comprises: Figure 4 As shown in the structure schematic diagram of the audio processing device provided by the embodiment of the application, the device comprises:

[0114] The tracking unit 410 is used for obtaining a video stream corresponding to a target area, and performing face tracking based on the video stream.

[0115] The first starting unit 420 is configured to, in a case where the tracking result of the face tracking includes a single first personnel identifier, track the single first personnel identifier in the video stream until the tracking of the single first personnel identifier fails, and start timing;

[0116] The second starting unit 430 is configured to, in a case where the tracking result of the face tracking includes a plurality of first personnel identifiers, track the plurality of first personnel identifiers in the video stream until the tracking of the plurality of first personnel identifiers all fails, and start timing;

[0117] The timing unit 440 is configured to continue timing until the timing duration reaches a preset duration or the tracking result changes to a second personnel identifier.

[0118] The segmentation unit 450 is configured to segment the audio stream corresponding to the target region based on the timing duration, or the first personnel identifier and the second personnel identifier.

[0119] The audio processing device provided by the embodiment of the present application realizes the automatic audio stream segmentation scheme based on personnel by performing face tracking on the video stream of the target region, and starting timing in a case where the tracking of the single first personnel identifier or all the first personnel identifiers fails, thereby overcoming the defects in the prior art that the audio can only be segmented by timing or manually operated, and thus the corresponding customers of the audio cannot be distinguished or the audio segmentation is not effective, and providing conditions for subsequently quickly and accurately finding a specified audio file according to actual needs.

[0120] Based on any of the above embodiments, the segmentation unit 450 includes:

[0121] The first segmentation sub-unit is configured to, in a case where the timing duration reaches the preset duration, segment the audio stream based on a first time, the first time being a time when the tracking of the single first personnel identifier or all the first personnel identifiers fails.

[0122] The second segmentation sub-unit is configured to, in a case where the tracking result changes to the second personnel identifier, segment the audio stream based on the first personnel identifier and the second personnel identifier.

[0123] Based on any of the above embodiments, the second segmentation sub-unit is specifically configured to:

[0124] In a case where there is a single first personnel identifier and the single first personnel identifier is consistent with the second personnel identifier, or in a case where there are a plurality of first personnel identifiers and at least one first personnel identifier in the plurality of first personnel identifiers is consistent with the second personnel identifier, return to perform face tracking;

[0125] In a case where the single first person identifier exists and is different from the second person identifier, or in a case where the multiple first person identifiers exist and are all different from the second person identifier, the audio stream is segmented based on the first time.

[0126] According to any one of the above embodiments, the device further comprises:

[0127] The acquisition unit is configured to acquire the multiple first audio segments segmented from the audio stream.

[0128] The merging unit is configured to merge the first audio segments with matching voiceprint features in the multiple first audio segments to obtain multiple second audio segments.

[0129] According to any one of the above embodiments, the merging unit is specifically configured to perform voiceprint feature matching on adjacent first audio segments in the multiple first audio segments, and merge the matching first audio segments.

[0130] According to any one of the above embodiments, the tracking unit is specifically configured to:

[0131] perform person detection on the video stream;

[0132] In a case where the detection result of the person detection indicates the presence of a person, perform face tracking on the video stream and collect the audio stream.

[0133] According to any one of the above embodiments, the preset time length is determined based on an interval time length between adjacent sample audio segments segmented from a sample audio stream.

[0134] According to any one of the above embodiments, Figure 5 As shown in FIG. 1, a structure schematic diagram of a recording terminal provided by an embodiment of the present application is shown, and as shown in FIG. 2, the recording terminal comprises: Figure 5 As shown in FIG. 1, a structure schematic diagram of a recording terminal provided by an embodiment of the present application is shown, and as shown in FIG. 2, the recording terminal comprises:

[0135] The microphone 510 is configured to collect an audio stream corresponding to the target area.

[0136] The camera 520 is configured to collect a video stream corresponding to the target area.

[0137] The processor 530 is connected with the microphone and the camera respectively, and is configured to acquire an audio stream and a video stream, perform face tracking based on the video stream, track the single first personnel identifier in the video stream in a case where the tracking result of the face tracking includes the single first personnel identifier until the tracking of the single first personnel identifier fails, start timing, or track the multiple first personnel identifiers in the video stream in a case where the tracking result of the face tracking includes the multiple first personnel identifiers until the tracking of all the first personnel identifiers fails, start timing, continue timing until the timing duration reaches a preset duration or the tracking result changes to a second personnel identifier, and split the audio stream based on the timing duration or the first personnel identifier and the second personnel identifier.

[0138] It should be noted that the microphone is an energy conversion device for converting a sound signal into an electric signal, and can realize audio acquisition. The microphone in the embodiment of the application can be a microphone array composed of multiple microphones to improve the effect of audio acquisition.

[0139] In the embodiment of the application, the recording terminal can record at a specified angle and suppress sound. During initialization, the effective sound pickup area is determined according to the pronunciation positions of the initialized customer and the clerk, and the target area can be set according to the determined effective sound pickup area. Figure 6 As shown in the scene schematic diagram of a customer handling business in a business hall provided by the embodiment of the application, the effective sound pickup area includes both the effective sound pickup area of the customer and the effective sound pickup area of the clerk, such as the dashed line area shown in FIG. Figure 6 .

[0140] Specifically, the recording terminal provided by the embodiment of the application can acquire audio and video of the customer at the counter or window of the business hall and acquire audio of the clerk. After determining the target area, the recording terminal detects whether a customer enters the target area through the camera. If so, the recording terminal starts recording the customer and the clerk through the microphone, acquires the audio stream corresponding to the target area, and acquires the video stream corresponding to the target area through the camera. The microphone and the camera are in communication connection with the processor, and can respectively transmit the acquired audio stream and video stream to the processor. The processor performs face tracking based on the acquired video stream, starts timing in a case where the tracking of the single first personnel identifier or all the first personnel identifiers fails, until the timing duration reaches a preset duration or the tracking result changes to a second personnel identifier, so as to split the audio stream based on the timing duration or the first personnel identifier and the second personnel identifier.

[0141] The audio terminal provided by the embodiment of the present application can track the face in the video stream of the target area, and start timing when tracking of the single first personnel identifier or all first personnel identifiers fails, so as to split the audio stream corresponding to the target area based on the timing duration, or based on the second personnel identifier obtained by changing the first personnel identifier and the tracking result, thereby realizing an automatic audio stream splitting scheme based on personnel, overcoming the defects in the prior art that the audio can only be split in time or manually, so that the corresponding customers of the audio cannot be distinguished and the audio splitting effect is not high, and providing conditions for quickly and accurately finding the specified audio file according to actual needs.

[0142] Other embodiments or specific implementation manners of the audio terminal provided by the embodiment of the present application can refer to the above-mentioned method embodiments, and will not be described here.

[0143] Figure 7 An example of a schematic diagram of the physical structure of an electronic device is shown in Figure 7 The electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 can communicate with each other through the communications bus 740. The processor 710 can invoke a logical instruction in the memory 730 to execute an audio processing method, which includes: obtaining a video stream corresponding to a target area, and tracking a face based on the video stream; in a case where the tracking result of the face tracking includes a single first personnel identifier, tracking the single first personnel identifier in the video stream until tracking of the single first personnel identifier fails, and starting timing; in a case where the tracking result of the face tracking includes a plurality of first personnel identifiers, tracking the plurality of first personnel identifiers in the video stream until tracking of all first personnel identifiers fails, and starting timing; continuously timing until the timing duration reaches a preset duration or the tracking result changes to a second personnel identifier; and splitting an audio stream corresponding to the target area based on the timing duration, or the first personnel identifier and the second personnel identifier.

[0144] In addition, the logic instructions in the memory 730 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0145] The present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements an audio processing method provided by each of the above methods, for example, including: obtaining a video stream corresponding to a target region, and performing face tracking based on the video stream; in a case where a tracking result of the face tracking includes a single first person identifier, tracking the single first person identifier in the video stream until tracking of the single first person identifier fails, and starting timing; in a case where the tracking result of the face tracking includes a plurality of first person identifiers, tracking the plurality of first person identifiers in the video stream until tracking of all the first person identifiers fails, and starting timing; continuing timing until a timing duration reaches a preset duration or the tracking result changes to a second person identifier; and based on the timing duration or the first person identifier and the second person identifier, splitting an audio stream corresponding to the target region.

[0146] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0147] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0148] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio processing method, characterized by, The method comprises the following steps: acquiring a video stream corresponding to a target area, and performing face tracking based on the video stream; in a case where the tracking result of the face tracking comprises a single first person identifier, tracking the single first person identifier in the video stream until the tracking of the single first person identifier fails, and starting timing; in a case where the tracking result of the face tracking comprises a plurality of first person identifiers, tracking the plurality of first person identifiers in the video stream until the tracking of all the first person identifiers fails, and starting timing; continuously timing until the timing duration reaches a preset duration or the tracking result changes to a second person identifier; in a case where the timing duration reaches the preset duration, splitting an audio stream corresponding to the target area based on a first time, the first time being a time when the tracking of the single first person identifier or all the first person identifiers fails; in a case where the tracking result changes to the second person identifier, splitting the audio stream corresponding to the target area based on the first person identifier and the second person identifier.

2. The audio processing method of claim 1, wherein, The splitting of the audio stream corresponding to the target area based on the first person identifier and the second person identifier comprises: in a case where there is a single first person identifier and the single first person identifier is consistent with the second person identifier, or in a case where there are a plurality of first person identifiers and at least one of the plurality of first person identifiers is consistent with the second person identifier, returning to perform face tracking; in a case where there is a single first person identifier and the single first person identifier is different from the second person identifier, or in a case where there are a plurality of first person identifiers and the plurality of first person identifiers are all different from the second person identifier, splitting the audio stream based on the first time.

3. The audio processing method of claim 1, wherein, The splitting of the audio stream corresponding to the target area further comprises: acquiring a plurality of first audio segments obtained by splitting; merging first audio segments with matching voiceprint features in the plurality of first audio segments to obtain a plurality of second audio segments.

4. The audio processing method of claim 3, wherein, The merging of the first audio segments with matching voiceprint features in the plurality of first audio segments comprises: performing voiceprint feature matching on adjacent first audio segments in the plurality of first audio segments, and merging the matching first audio segments.

5. The audio processing method of any one of claims 1-4, wherein, The face tracking based on the video stream comprises: performing person detection on the video stream; in a case where the detection result of the person detection indicates the presence of a person, performing face tracking on the video stream and collecting the audio stream.

6. The audio processing method of any one of claims 1-4, wherein: The preset duration is determined based on an interval duration between adjacent sample audio segments obtained by splitting a sample audio stream.

7. An audio processing apparatus, characterized by comprising: The method comprises the following steps: a tracking unit configured to acquire a video stream corresponding to a target area, and perform face tracking based on the video stream; a first starting unit configured to, in a case where the tracking result of the face tracking comprises a single first person identifier, track the single first person identifier in the video stream until the tracking of the single first person identifier fails, and start timing; The second starting unit is configured to, in a case where the tracking result of the face tracking includes a plurality of first personnel identifiers, track the plurality of first personnel identifiers in the video stream until tracking of the plurality of first personnel identifiers all fails, and start timing; The timing unit is configured to continuously time until a timing duration reaches a preset duration or the tracking result changes to a second personnel identifier; The cutting unit is configured to, in a case where the timing duration reaches the preset duration, cut the audio stream corresponding to the target region based on a first time, the first time being a time when tracking of the single first personnel identifier or all first personnel identifiers fails; and in a case where the tracking result changes to a second personnel identifier, cut the audio stream corresponding to the target region based on the first personnel identifier and the second personnel identifier.

8. A recording terminal, characterized by comprising: The microphone is configured to collect an audio stream corresponding to a target region. The camera is configured to collect a video stream corresponding to the target region. The processor is connected with the microphone and the camera respectively, and is configured to acquire the audio stream and the video stream, and perform face tracking based on the video stream. In a case where the tracking result of the face tracking includes a single first personnel identifier, the single first personnel identifier is tracked in the video stream until tracking of the single first personnel identifier fails, timing is started, or, in a case where the tracking result of the face tracking includes a plurality of first personnel identifiers, the plurality of first personnel identifiers are tracked in the video stream until tracking of all first personnel identifiers all fails, timing is started; the timing is continuously performed until a timing duration reaches a preset duration or the tracking result changes to a second personnel identifier; and in a case where the timing duration reaches the preset duration, the audio stream is cut based on a first time, the first time being a time when tracking of the single first personnel identifier or all first personnel identifiers fails, and in a case where the tracking result changes to a second personnel identifier, the audio stream is cut based on the first personnel identifier and the second personnel identifier. The processor implements the audio processing method according to any one of claims 1 to 6 when executing the program.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program implements the audio processing method according to any one of claims 1 to 6 when executed by the processor. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Video segmentation method, system and device, and storage medium

    CN112565885A

  • Audio and video control method with wireless microphone intelligent tracking function

    CN115695708A