Multi-role collaborative customer service agent method
By establishing an audio-visual frame window mapping table and a phase perception model, the accuracy and efficiency issues of video authenticity verification in customer service intelligent agents were resolved. This enabled efficient video type identification and anomaly localization, reduced labor costs, and improved the anti-fraud capabilities of remote business processing.
Patent Information
- Application Number
- CN202511453218.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing customer service AI agents suffer from insufficient accuracy and low efficiency in verifying the authenticity of videos, making it difficult to effectively prevent fraud risks. In particular, when lighting and audio equipment are affected by power frequency, single-modal data features are difficult to identify spliced or replayed videos, and the lack of pre-quality control leads to misjudgments and high labor costs.
By establishing an audio-visual frame window mapping table, extracting the power frequency phase trajectory and calculating the signal-to-noise ratio, and utilizing linear scale mapping and phase perception models, combined with sliding window local anomaly localization, accurate identification of audio-visual cross-modal correlation and breakpoint localization can be achieved, reducing the cost of manual review.
It improves the accuracy and efficiency of video authenticity verification, reduces labor costs, can quickly support business decisions, effectively prevent business fraud risks, and improve the efficiency of remote business processing.
Smart Images

Figure CN120911508A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, more specifically, it relates to a multi-role cooperative customer service intelligent agent method. BACKGROUND
[0002] With the popularization of artificial intelligence in the field of customer service, remote business handling usually relies on video interaction containing audio, such as online account opening, fault repair, etc., and video authenticity becomes the key to risk control; among them, the video that is forged, spliced or played back may cause identity misjudgment, business fraud and other problems, therefore, the authenticity of the video needs to be strictly verified.
[0003] The existing video verification method of the customer service intelligent agent has the following obvious limitations: 1. Usually relying on single modal data features, it is difficult to utilize the common physical correlation of audio and video, in actual scenes, lighting devices will produce periodic brightness fluctuations due to the influence of power grid frequency, and audio devices will also produce associated fluctuations due to the influence of power frequency radiation, the audio and video power frequency features of the real video should have a stable corresponding relationship, and the spliced or played back video will destroy this correlation, making it difficult to accurately identify the abnormal type of the video; 2. Lack of quality pre-control, when the video brightness fluctuation is weak and the audio noise is strong, the power frequency features are masked, and direct analysis may cause misjudgment; 3. Low local anomaly positioning efficiency, even if the overall video anomaly is found, manual frame-by-frame inspection of splicing or editing breakpoints is required, and in the case of large amount of customer service scene video, the manual cost is high and time-consuming, which cannot quickly support business decision-making.
[0004] The above problems lead to insufficient accuracy and low efficiency of the customer service intelligent agent in video authenticity verification, which not only affects the business handling speed, but also makes it difficult to effectively prevent fraud risks, becoming a key bottleneck for the intelligent upgrading of customer service scenes. SUMMARY
[0005] The present application provides a multi-role cooperative customer service intelligent agent method to solve the technical problems in the background art.
[0006] The present application provides a multi-role cooperative customer service intelligent agent method, comprising the following steps: Step S101, obtaining a target video containing audio and its metadata, and establishing a frame window set based on video frames, so that each video frame is associated with the audio samples in its time interval to construct a frame window mapping table; The metadata includes: video frame rate, audio sampling rate, total number of video frames, frame-by-frame timestamp sequence and total number of audio samples; Step S102, performing feature extraction on each video frame to obtain a video phase trajectory; Step S103, traversing the frame window mapping table, feature extraction is performed on the audio samples in the time interval of each video frame to obtain an audio phase trajectory, and a video side signal-to-noise ratio and an audio side signal-to-noise ratio are respectively calculated; Step S104, linear scale mapping of the audio phase trajectory to the video phase trajectory is established to obtain a linear scale parameter, and a global consistency score is calculated according to the linear scale parameter; The linear scale parameter includes a scale parameter and an offset parameter; Step S105, a feature sequence is constructed based on the video phase trajectory and the audio phase trajectory, and input into the phase perception model after training, and the type of the target video is output; The type of the target video includes normal, splicing, playback and screen shooting; Step S106, a local consistency score is calculated through a sliding window, and if the local consistency score is less than a third threshold value, the local consistency score is marked as a suspicious window, and a breakpoint index in the suspicious window is determined according to a phase difference change rate.
[0007] Further, the time interval of each video frame is determined as a frame window, and the timestamp of each audio sample is determined, the frame index of each video frame is taken as a key, and the audio sample whose timestamp falls within the frame window corresponding to each video frame is taken as a value to construct a frame window mapping table; The time interval of the kth video frame is represented as: , wherein 1≤k≤K, K represents the total number of video frames, represents the starting time of the kth video frame in the frame timestamp sequence, represents the video frame rate; The timestamp of the nth audio sample is represented as: , wherein 1≤n≤N, N represents the total number of audio samples, represents the audio sampling rate.
[0008] Further, feature extraction is performed on each video frame to obtain a video phase trajectory, including the following steps: Step S201, each video frame is subjected to grayscale processing to obtain a grayscale matrix; Step S202, the average value of all pixel gray values of each grayscale matrix is calculated as a brightness value, the average value of the brightness values of all grayscale matrices is calculated as a global brightness value, and each grayscale matrix is subtracted from the global brightness value to construct a brightness sequence; Step S203, a band-pass filter is applied to the brightness sequence to retain frequency components related to the power frequency, and a band-pass brightness sequence is obtained; Step S204, the band-pass brightness sequence is converted into an analytic signal through Hilbert transform, the principal value phase is extracted and the periodic jump is eliminated, and a video phase trajectory is obtained; Each element value of the video phase trajectory is represented by a real scalar, corresponding to the video phase state of a video frame.
[0009] Further, a band-pass filter is applied to the audio sample sequence in the time interval in which each video frame is located, the frequency components related to the power frequency are retained, and the band-pass audio sequence is obtained by taking the average value of the filtered audio samples in each frame window and aggregating the frame window. The band-pass audio sequence is converted into an analytic signal by Hilbert transform, the principal value phase is extracted and the periodic jump is eliminated, and the audio phase trajectory is obtained. The total number of elements of the audio phase trajectory is the same as that of the video phase trajectory. Each element value of the audio phase trajectory is represented by a real scalar, corresponding to the audio phase state of a video frame. The Fourier transform is performed on the band-pass luminance sequence and the band-pass audio sequence respectively to obtain the corresponding energy spectrum density. The energy spectrum density of the main peak frequency point and its neighborhood is integrated to obtain the video side main peak power and the audio side main peak power. The energy spectrum density corresponding to the band-pass luminance sequence and the band-pass audio sequence is integrated to obtain the video passband total power and the audio passband total power. The video passband total power is subtracted by the video side main peak power to obtain the video side residual power. The audio passband total power is subtracted by the audio side main peak power to obtain the audio side residual power. The video side signal-to-noise ratio and the audio side signal-to-noise ratio are calculated respectively.
[0010] Further, the minimum value of the video side signal-to-noise ratio and the audio side signal-to-noise ratio is taken, and if the value is less than a first threshold value, it is judged to enter the artificial review queue, wherein the first threshold value is a self-defined parameter.
[0011] Further, the video phase trajectory and the audio phase trajectory are combined to construct a linear equation, and the proportional parameter and the offset parameter of the linear equation are obtained by least square closed-form solution. The calculation formula of the linear equation is as follows: wherein 1≤i≤L, L represents the total number of elements of the video phase trajectory or the audio phase trajectory, and represent the i-th element value of the video phase trajectory and the audio phase trajectory respectively, a represents the proportional parameter, and b represents the offset parameter.
[0012] Further, the residual sequence is calculated according to the video phase trajectory, the audio phase trajectory, the proportional parameter and the offset parameter, and the global consistency score is calculated according to the residual sequence. If the global consistency score is greater than or equal to a second threshold value, the global consistency is output, otherwise the global inconsistency is output, wherein the second threshold value is a self-defined parameter. The calculation formula of the i-th element value of the residual sequence is as follows: Where 1≤i≤L, and L represents the total number of elements in the video phase trajectory or audio phase trajectory. and Let a and b represent the i-th element values of the video phase trajectory and the audio phase trajectory, respectively, where a represents the scaling parameter and b represents the offset parameter. Global Consistency Score The calculation formula is as follows: ,in The sample variance of the residual sequence. This represents the sample variance of the video phase trajectory.
[0013] Furthermore, the number of sequence units in the feature sequence is the same as the total number of elements in the video phase trajectory and the audio phase trajectory. The sequence units of the feature sequence are represented by feature vectors, which consist of element values of the video phase trajectory, element values of the audio phase trajectory, circular phase difference, video phase increment, audio phase increment, frequency locking error, video phase coding, and audio phase coding. The circular phase difference of the eigenvector corresponding to the i-th sequence unit The calculation formula is as follows: ,in This represents the value of the i-th element in the residual sequence. Indicates rounding down; The video phase increment and the audio phase increment represent the difference between the current element value and the previous element value of the video phase trajectory and the audio phase trajectory, respectively. The frequency locking error is equal to the video phase increment minus the product of the scaling parameter and the audio phase increment; Video phase coding of the feature vector corresponding to the i-th sequence unit The calculation formula is as follows: ,in Indicates the signal-to-noise ratio on the video side. This represents the value of the i-th element of the video phase trajectory. This represents the video phase adjustment parameters, which are user-defined parameters. The audio phase encoding of the feature vector corresponding to the i-th sequence unit The calculation formula is as follows: ,in Indicates the signal-to-noise ratio on the audio side. This represents the value of the i-th element of the audio phase trajectory. This represents the audio phase adjustment parameter, which is a user-defined parameter.
[0014] Further, the phase-aware model is constructed based on a gated neural network model, and a product of a weight parameter and a concentration degree is added in a calculation formula of a reset gate and an update gate. The calculation formula of the concentration degree F is as follows: , wherein denotes a frequency-locked error, and denote a first weight parameter and a second weight parameter, respectively; A sample label of a training sample for training the phase-aware model is obtained by manual labeling.
[0015] Further, a size of the sliding window is a self-defined parameter, calculation formulas of the local consistency score and the global consistency score are the same, a third threshold value is a self-defined parameter, a frame index of a video frame corresponding to a maximum change rate of element values of a residual sequence corresponding to the sliding window is selected as the breakpoint index by calculating the change rate.
[0016] The present application has the advantages that: the present application establishes audio-video time association based on video frames, extracts audio-video power frequency phase trajectories and calculates signal-to-noise ratios, first, low-quality data is shunted to manual review through quality gating to avoid misjudgment diffusion; then, the audio-video cross-modal correlation degree is accurately judged through linear scale mapping and global consistency scoring, and the phase-aware model is combined to realize accurate identification of video types, and a sliding window is used to locate local abnormal breakpoints for the convenience of manual rapid review; the present application uses power frequency physical characteristics to ensure the accuracy of verification, reduces the cost of manual work through quality pre-control, intelligent type identification and breakpoint positioning, and the consistent conclusion and abnormal evidence output can support multi-role closed-loop cooperation, effectively prevent business fraud risks, improve the efficiency of remote business handling in customer service scenarios, and break through the bottleneck of low efficiency of traditional single modal verification and manual review. BRIEF DESCRIPTION OF DRAWINGS
[0017] Fig. 1 is a flowchart of a multi-role cooperative customer service agent method of the present application; Fig. 2 is a flowchart of feature extraction of each video frame to obtain a video phase trajectory of the present application; Fig. 3 is a test result graph of the phase-aware model of the present application. DETAILED DESCRIPTION
[0018] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the implementations discussed are merely for illustration and that the elements of the discussions can be modified, supplemented, or omitted in different examples. Additionally, features described in relation to some examples can also be combined in other examples.
[0019] It should be noted that the technical terms or scientific terms used in one or more embodiments of the present application should be understood as the general meaning understood by a person having ordinary skill in the art to which the present application pertains, unless otherwise defined. The terms "first", "second", and the like used in one or more embodiments of the present application do not represent any order, number, or importance, but are used to distinguish different components. The terms "include" or "contain" and the like mean that the elements or objects appearing before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connected" or "linked" and the like do not mean only physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are used only to indicate relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly.
[0020] As shown in Figs. 1-3 A multi-role collaborative customer service agent method includes the following steps: Step S101, obtaining a target video containing audio and metadata thereof, and establishing a frame window set based on video frames, so that each video frame is associated with audio samples in its time interval to construct a frame window mapping table; The metadata includes: video frame rate, audio sampling rate, total number of video frames, frame-by-frame timestamp sequence, and total number of audio samples; Step S102, feature extraction is performed on each video frame to obtain a video phase trajectory; Step S103, traversing the frame window mapping table, feature extraction is performed on the audio samples in the time interval of each video frame to obtain an audio phase trajectory, and video side signal-to-noise ratio and audio side signal-to-noise ratio are calculated and obtained, respectively; Step S104, linear scale mapping of the audio phase trajectory to the video phase trajectory is established to obtain linear scale parameters, and a global consistency score is calculated therefrom; The linear scale parameters include: a scale parameter and an offset parameter; Step S105, a feature sequence is constructed based on the video phase trajectory and the audio phase trajectory, input into a phase-aware model after training, and the type of the target video is output; The type of the target video includes normal, splicing, playback and screen capture. In step S106, the local consistency score is calculated by a sliding window. If the local consistency score is less than a third threshold, the window is marked as a suspicious window, and a breakpoint index in the suspicious window is determined according to a phase difference change rate.
[0021] In an embodiment of the present application, if the target video does not satisfy the monotonicity check and the full-time domain coverage check, it is determined that the target video is incorrect, and the user is prompted to upload the video again. The monotonicity check means that each timestamp in the sequence of frame-by-frame timestamps must be sequentially increased. The full-time domain coverage check means that the time interval of all video frames must completely cover the total duration of the target video.
[0022] It should be noted that the video frame rate represents the number of picture frames contained in a video per second. For example, if the video frame rate is 30 frames per second, it means that the video contains 30 consecutive pictures per second. The audio sampling rate represents the number of sampling points contained in the audio per second. The total number of video frames represents the total number of pictures contained in the video. For example, if the total duration of the video is 10 seconds and the video frame rate is as described above, the corresponding total number of video frames is 300 frames. The sequence of frame-by-frame timestamps represents the starting time of each video frame. For example, the 0th frame is 0 seconds, the 1st frame is 1 / 30 seconds, and so on. The total number of audio samples is equal to the total duration of the video multiplied by the audio sampling rate. For example, if the audio sampling rate is 44100 samples per second (44.1 kHz) and the total duration of the video is as described above, the total number of audio samples is 441000.
[0023] In an embodiment of the present application, the time interval of each video frame is first determined as a frame window, and then the timestamp of each audio sample is determined. The frame index of each video frame is taken as a key, and the audio sample whose timestamp falls within the frame window corresponding to each video frame is taken as a value to construct a frame window mapping table. The time interval of the kth video frame is represented as: wherein 1≤k≤K, K represents the total number of video frames, represents the starting time of the kth video frame in the sequence of frame-by-frame timestamps, represents the video frame rate. The timestamp of the nth audio sample is represented as: wherein 1≤n≤N, N represents the total number of audio samples, represents the audio sampling rate.
[0024] It should be noted that determining the time interval of each video frame can ensure that the time intervals of adjacent frames have no overlap, avoid the same audio sample being repeatedly associated with multiple frames, and enable each video frame to uniquely correspond to the audio sample within its time interval, thereby providing a strict time reference for subsequent cross-modal analysis.
[0025] In an embodiment of the present application, as Fig. 2As shown, feature extraction is performed on each video frame to obtain a video phase trajectory, including the following steps: Step S201, each video frame is subjected to grayscale processing to obtain a grayscale matrix; Step S202, the average value of the gray values of all pixels of each grayscale matrix is calculated as a brightness value, the average value of the brightness values of all grayscale matrices is calculated as a global brightness value, and the brightness value corresponding to each grayscale matrix is subtracted from the global brightness value to construct a brightness sequence; Step S203, a band-pass filter is applied to the brightness sequence to retain the frequency components related to the power frequency, and a band-pass brightness sequence is obtained; The band-pass filter can be a FIR (Finite Impulse Response) band-pass filter or an IIR (Infinite Impulse Response) band-pass filter, which will not be described here; Step S204, the band-pass brightness sequence is converted into an analytic signal through Hilbert transform, the principal value phase is extracted and the periodic jump is eliminated, and a video phase trajectory is obtained; Each element value of the video phase trajectory is represented by a real scalar, corresponding to the video phase state of a video frame.
[0026] It should be noted that the global brightness value is used to represent the average light and dark level of the target video, and the brightness value of each video frame is subtracted from the global brightness value to eliminate the influence of the overall illumination level on the analysis, thereby highlighting the fluctuations of the brightness of each video frame relative to the global average level. The power frequency refers to the standard frequency of the power system, which is usually 50Hz or 60Hz. Influenced by the power frequency, the lighting device will produce periodic brightness fluctuations related to the power frequency, and the frequency may be the power frequency itself or its harmonics such as 100Hz, 120Hz, etc. The band-pass filter retains the frequency components related to the power frequency, effectively filtering out low-frequency disturbances such as environmental light changes and high-frequency disturbances such as sensor noise, and highlighting the periodic fluctuation characteristics of the brightness sequence caused by the power frequency. The real part of the analytic signal is the original band-pass brightness sequence, and the imaginary part is the orthogonal component obtained by Hilbert transform of the sequence. The principal value phase is the angle between the analytic signal and the positive direction of the real axis in the complex plane, and its value range is between -π and +π, which is used to represent the phase state of the signal at that time. When the signal phase continuously changes over time and exceeds +π, a periodic jump occurs, i.e. from +π to -π. Through phase unwrapping processing, such jumps can be eliminated, i.e. when a jump is detected, the phase value is corrected by accumulating 2π, for example, corrected to +π+2π=3π, so that the phase changes continuously and smoothly over time.
[0027] In one embodiment of the present application, a band-pass filter is applied to the sequence of audio samples in the time interval in which each video frame is located, the frequency components related to the power frequency are retained, and a band-pass audio sequence is obtained by taking the average value of the filtered audio samples in each frame window and aggregating the frame windows. The band-pass audio sequence is converted into an analytic signal by Hilbert transform, the principal value phase is extracted and periodic jumps are eliminated to obtain an audio phase trajectory. The total number of elements of the audio phase trajectory is the same as the total number of elements of the video phase trajectory. Each element value of the audio phase trajectory is represented by a real scalar, which corresponds to the audio phase state of a video frame. The Fourier transform is performed on the band-pass luminance sequence and the band-pass audio sequence respectively to obtain the corresponding energy spectrum density. The energy spectrum density of the main peak frequency point and its neighborhood is integrated to obtain the video side main peak power and the audio side main peak power. The energy spectrum density corresponding to the band-pass luminance sequence and the band-pass audio sequence is integrated to obtain the video passband total power and the audio passband total power. The video passband total power is subtracted by the video side main peak power to obtain the video side residual power. The audio passband total power is subtracted by the audio side main peak power to obtain the audio side residual power. The video side signal-to-noise ratio and the audio side signal-to-noise ratio are calculated respectively. The video side signal-to-noise ratio The calculation formula is as follows: , wherein and represent the video side main peak power and the video side residual power respectively. The audio side signal-to-noise ratio The calculation formula is as follows: , wherein and represent the audio side main peak power and the audio side residual power respectively.
[0028] It should be noted that the neighborhood width is a self-defined parameter, and preferably, the neighborhood width is set to ±5Hz. The audio sample represents the numerical record of the audio signal at the discrete time point, i.e. the instantaneous amplitude of the audio signal. According to the total number of video frames and the total number of audio samples, the number of audio samples contained in each frame window is 1470, and the total number of elements of the audio phase trajectory and the total number of elements of the video phase trajectory are both 300. The video side signal-to-noise ratio is used to measure the quality of the luminance signal related to the power frequency in the target video, and the audio side signal-to-noise ratio is used to measure the quality of the audio signal related to the power frequency in the target video. The larger the value of the signal-to-noise ratio, the less the interference.
[0029] In one embodiment of the present application, the minimum value of the video side signal-to-noise ratio and the audio side signal-to-noise ratio is taken, and if the value is less than a first threshold value, it is judged to enter the artificial review queue, wherein the first threshold value is a self-defined parameter, and preferably, the first threshold value is set to 10dB.
[0030] It should be noted that when the smaller signal-to-noise ratio is less than the first threshold value, it indicates that the power frequency related signal in the video or audio is too weak and the noise is too much, and the subsequent analysis is prone to error; therefore, it is directly entered into the manual review queue, in order to avoid misjudgment through manual inspection, and to ensure that the final judgment result is more real and reliable; the value range of the first threshold value is usually set to be between 8 to 15 dB, for example, when the signal-to-noise ratio is less than 8 dB, it indicates that the video or audio signal is almost submerged by noise, and it is difficult to extract effective information, and when the signal-to-noise ratio is higher than 15 dB, it indicates that the video or audio quality is good, but the requirement may be too strict, which will exclude some samples that can still be analyzed although there is slight noise, therefore, the first threshold value is set to 10 dB to balance the signal quality and practicability.
[0031] In an embodiment of the present application, the video phase trajectory and the audio phase trajectory are combined to construct a linear equation, the proportional parameter and the offset parameter of the linear equation are obtained by least square closed-form solution, and the calculation formula of the linear equation is as follows: , wherein 1≤i≤L, L represents the total number of elements of the video phase trajectory or the audio phase trajectory, that is, L=K, K represents the total number of video frames, and respectively represent the i-th element value of the video phase trajectory and the audio phase trajectory, a represents the proportional parameter, and b represents the offset parameter; The calculation formula of the proportional parameter a by least square closed-form solution is as follows: , wherein represents the sample covariance between the audio phase trajectory and the video phase trajectory, represents the sample variance of the audio phase trajectory; The calculation formula of the offset parameter b by least square closed-form solution is as follows: , wherein and respectively represent the sample mean of the video phase trajectory and the audio phase trajectory.
[0032] It should be noted that the phase trajectories of the video and the audio are both affected by the power frequency, but due to the difference of the collection devices (for example, the camera collects the video and the microphone collects the audio), there is usually a stable linear relationship between the two, and the linear equation is constructed by combining the two, which can eliminate the influence of the device difference, so that the two can be compared under the same standard.
[0033] In an embodiment of the present application, the residual sequence is calculated according to the video phase trajectory, the audio phase trajectory, the proportional parameter and the offset parameter, the global consistency score is calculated according to the residual sequence, and if the global consistency score is greater than or equal to the second threshold value, the global consistency is output, otherwise the global inconsistency is output. the i-th element value of the residual sequence The calculation formula of is as follows: , wherein 1≤i≤L, L represents the total number of elements of the video phase track or the audio phase track, and respectively represent the i-th element value of the video phase track and the audio phase track, a represents a proportional parameter, and b represents an offset parameter; global consistency score The calculation formula of is as follows: , wherein represents the sample variance of the residual sequence, represents the sample variance of the video phase track; The second threshold value is a user-defined parameter, and preferably, the second threshold value is set to 0.8.
[0034] It should be noted that the residual sequence is constructed to measure the fitting effect of the linear equation, and the smaller the residual value is, the better the linear equation can reflect the corresponding relationship between the video and the audio, that is, the video and the audio phase change are consistent; the sample variance of the residual sequence represents the deviation of the linear equation fitting, and the sample variance of the video phase track represents the total fluctuation of the video phase itself, the smaller the ratio of the two is, the closer the global consistency score is to 1, and the higher the overall consistency of the video and the audio phase track is, and the stronger the corresponding relationship is; in addition, the video and the audio that are edited and spliced have a lower global consistency score due to different signal sources, and the value range of the second threshold value is usually set to be between 0.7 and 0.95, which can effectively distinguish whether the video and the audio are truly synchronized and whether there is human editing.
[0035] In an embodiment of the present application, the number of sequence units of the feature sequence is the same as the total number of elements of the video phase track and the audio phase track, the sequence unit of the feature sequence is represented by a feature vector, and the feature vector is composed of the element value of the video phase track, the element value of the audio phase track, the circular domain phase difference, the video phase increment, the audio phase increment, the frequency lock error, the video phase encoding, and the audio phase encoding, that is, the feature vector has a total of 8 dimension values; The circular domain phase difference of the i-th sequence unit corresponding to the feature vector The calculation formula of is as follows: , wherein represents the i-th element value of the residual sequence, represents a down rounding; The video phase increment and the audio phase increment respectively represent the difference between the current element value and the previous element value of the video phase track and the audio phase track; The frequency locking error is equal to the video phase increment minus the product of the scaling parameter and the audio phase increment; Video phase coding of the feature vector corresponding to the i-th sequence unit The calculation formula is as follows: ,in Indicates the signal-to-noise ratio on the video side. This represents the value of the i-th element of the video phase trajectory. This indicates the video phase adjustment parameter, which is a custom parameter. Preferably, the video phase adjustment parameter is set to 120 (1 / rad). The audio phase encoding of the feature vector corresponding to the i-th sequence unit The calculation formula is as follows: ,in Indicates the signal-to-noise ratio on the audio side. This represents the value of the i-th element of the audio phase trajectory. This represents the audio phase adjustment parameter, which is a custom parameter. Preferably, the audio phase adjustment parameter is set to 160 (1 / rad).
[0036] It should be noted that increasing the dimensionality of the original phase trajectory can preserve the periodic structure and linear trend of the phase, avoid discontinuities caused by phase jumps, and improve the expressive ability of phase-locked and frequency-locked modes, providing a learnable basis for the phase-aware model and enhancing the robustness of the model. In addition to adjusting the dimensions, the video phase adjustment parameters and audio phase adjustment parameters are also used to control the phase encoding output to remain within a stable numerical range. That is, when the phase amplitude is small or the signal-to-noise ratio is high, the exponential term is prone to surge, while maintaining a stable numerical range can help accelerate model convergence.
[0037] In one embodiment of the present invention, the phase sensing model is constructed based on a gated neural network model (GRU), and the product of weight parameters and concentration is added to the calculation formulas of the reset gate and the update gate. The formula for calculating the concentration degree F is as follows: ,in Indicates frequency locking error. and These represent the first weight parameter and the second weight parameter, respectively. The sample labels for the training samples used to train the phase-sensing model were obtained through manual annotation.
[0038] Specifically, the number of reset gates and update gates is the same as the number of sequence units in the feature sequence, and they correspond one-to-one.
[0039] Specifically, the t-th reset door The calculation formula is as follows: wherein denotes the t-th concentration, denotes the hidden vector of the t-1-th reset gate output (dimension number is a custom parameter, for example, set to 64), is assigned to 0, denotes the feature vector corresponding to the t-th sequence unit of the feature sequence of the t-th reset gate input, and denote the first weight parameter and the second weight parameter of the t-th reset gate, respectively, denotes the bias parameter of the t-th reset gate, denotes the Sigmoid activation function; Specifically, the t-th update gate is calculated as follows: wherein and denote the first weight parameter and the second weight parameter of the t-th update gate, respectively, denotes the bias parameter of the t-th update gate.
[0040] It should be noted that the weight parameters and the bias parameters in the phase-aware model are all learnable parameters, and the cross-entropy function is selected as the loss function, and the weight parameters and the bias parameters in the phase-aware model are updated in reverse through the gradient optimizer (such as Adam, RMSProp, etc.) to minimize the loss value calculated by the loss function, until the maximum iteration number is reached or the loss value calculated by the loss function is within the set range, then the training of the model is completed; in addition, the phase-aware model can also be constructed based on the LSTM (Long Short-Term Memory Recurrent Neural Network Model) and the Transfomer (Transformer) model, which will not be described here.
[0041] It should be noted that the phase concentration is introduced in the reset gate and the update gate, which directly converts the phase locking strength into an adaptive weight. High concentration indicates that the phase difference of the video frame is stable in the circular domain, the gate is more open, which encourages the network to retain and propagate such reliable information. Low concentration indicates that the phase locking is disturbed or there are editing traces, and the gate tends to be closed, thereby suppressing the interference of noise frames on the state, embedding the priori of physical quantities into the time sequence memory, which can significantly improve the gradient signal-to-noise ratio and the convergence speed, reduce the sensitivity of the model to hyperparameters, and thus enhance the robustness of the model.
[0042] In one embodiment of the present application, the size of the sliding window is a self-defined parameter, preferably, the size of the sliding window is set to 1 / 10 of the total number of frames of the video, the calculation formula of the local consistency score and the global consistency score is the same, and the third threshold value is a self-defined parameter, preferably, the third threshold value is set to 0.75. By calculating the element value change rate of the residual sequence corresponding to the sliding window, the frame index of the video frame corresponding to the maximum change rate is selected as the breakpoint index.
[0043] It should be noted that the breakpoint index indicates that the corresponding video frame is likely to have splicing, playback and screen shooting. In addition, multiple frame indexes of video frames with large change rates can be selected for manual review.
[0044] In one embodiment of the present application, after the calculation of the phase trajectory and the signal-to-noise ratio is completed, the system enters a parallel branch: one calculates the global consistency score and outputs whether it meets the second threshold value; the second reviews and identifies the type of the target video through the phase perception model, and outputs the video type; the third calculates the local consistency score through the sliding window and locates the breakpoint, and outputs the suspicious window and the breakpoint index; only when all the threshold values are met and the phase perception model does not find any abnormalities, a consistent conclusion in the presence is given, and the technical support continues to process, ensuring that multiple roles cooperate based on the same evidence closed loop.
[0045] As shown in Fig. 3 The phase perception model after training is tested to obtain a confusion matrix, and the number of test samples of each video type is 100, wherein the test accuracy rates of the video types of normal, splicing, playback and screen shooting are 96%, 96%, 95% and 97%, respectively.
[0046] It should be noted that the interval and the threshold value are set for easy comparison, and the size of the threshold value depends on the number of sample data and the base number set by the person skilled in the art for each group of sample data, as long as it does not affect the proportional relationship of the parameters and the quantized values. And the above formula is a dimensionless calculation of the value, and the formula is obtained by software simulation of a large amount of data to obtain a formula of the nearest real situation, and the preset parameters in the formula are set by the person skilled in the art according to the actual situation.
[0047] The above describes the embodiments of the present application, but the present application is not limited to the above specific embodiments, and the above specific embodiments are only illustrative and not limiting, and those skilled in the art can make many forms under the inspiration of the present application, which are all within the protection scope of the present application.
Claims
1. A multi-role collaborative customer service agent method, characterized in that, The method comprises the following steps: Step S101, obtaining a target video containing audio and metadata thereof, and establishing a frame window set based on video frames, so that each video frame is associated with an audio sample in a time interval thereof to construct a frame window mapping table; The metadata comprises a video frame rate, an audio sampling rate, a total number of video frames, a timestamp sequence of each frame, and a total number of audio samples; Step S102, performing feature extraction on each video frame to obtain a video phase trajectory; Step S103, traversing the frame window mapping table to perform feature extraction on the audio samples in the time interval of each video frame to obtain an audio phase trajectory, and respectively calculating a video side signal-to-noise ratio and an audio side signal-to-noise ratio; Step S104, establishing a linear scale mapping of the audio phase trajectory to the video phase trajectory to obtain a linear scale parameter, and calculating a global consistency score according to the linear scale parameter; The linear scale parameter comprises a scale parameter and an offset parameter; Step S105, constructing a feature sequence based on the video phase trajectory and the audio phase trajectory, inputting the feature sequence into a phase perception model trained to be completed, and outputting a type of the target video; The type of the target video comprises normal, splicing, playback, and screen capture; Step S106, calculating a local consistency score through a sliding window, determining whether the local consistency score is less than a third threshold value, and if so, marking the sliding window as a suspicious window, and determining a breakpoint index in the suspicious window according to a phase difference change rate. 2.The multi-role collaborative customer service agent method of claim 1, wherein, First, the time interval of each video frame is determined as a frame window, and then the timestamp of each audio sample is determined, each video frame index is taken as a key, and the audio samples whose timestamps fall within the frame window corresponding to each video frame are taken as values to construct a frame window mapping table; The time interval of the kth video frame is represented as: where 1≤k≤K, K represents the total number of video frames, represents the starting time of the kth video frame in the sequence of frame-by-frame timestamps, represents the video frame rate; the timestamp of the nth audio sample is represented as: where 1≤n≤N, N represents the total number of audio samples, represents the audio sampling rate.
3. The multi-role collaborative customer service agent method of claim 1, wherein, The feature extraction on each video frame to obtain a video phase trajectory comprises the following steps: Step S201, performing grayscale processing on each video frame to obtain a grayscale matrix; Step S202, calculating the average value of all pixel grayscale values of each grayscale matrix as a luminance value, calculating the average value of the luminance values of all grayscale matrices as a global luminance value, and constructing a luminance sequence by subtracting the global luminance value from the luminance value corresponding to each grayscale matrix; Step S203, applying a band-pass filter to the luminance sequence to retain frequency components related to the power frequency, and obtaining a band-pass luminance sequence; Step S204, converting the band-pass luminance sequence into an analytic signal through Hilbert transform, extracting the principal value phase thereof, and eliminating periodic jumps to obtain a video phase trajectory; Each element value of the video phase trajectory is represented by a real scalar, corresponding to the video phase state of a video frame.
4. The multi-role collaborative customer service agent method of claim 3, wherein, A band-pass filter is applied to the audio sample sequence in the time interval of each video frame to retain frequency components related to the power frequency, and a band-pass audio sequence is obtained by taking the average value of the filtered audio samples in each frame window for frame window aggregation, the band-pass audio sequence is converted into an analytic signal through Hilbert transform, the principal value phase thereof is extracted, and periodic jumps are eliminated to obtain an audio phase trajectory, the total number of element values of the audio phase trajectory is the same as that of the video phase trajectory, and each element value of the audio phase trajectory is represented by a real scalar, corresponding to the audio phase state of a video frame; The Fourier transform is performed on the band-pass luminance sequence and the band-pass audio sequence respectively to obtain corresponding energy spectrum densities, and the energy spectrum densities of the main peak frequency point and its neighborhood are integrated to obtain the video-side main peak power and the audio-side main peak power, respectively; the energy spectrum densities corresponding to the band-pass luminance sequence and the band-pass audio sequence are integrated to obtain the video passband total power and the audio passband total power, respectively; the video passband total power is subtracted by the video-side main peak power to obtain the video-side residual power, and the audio passband total power is subtracted by the audio-side main peak power to obtain the audio-side residual power; and the video-side signal-to-noise ratio and the audio-side signal-to-noise ratio are calculated respectively.
5. The multi-role collaborative customer service agent method of claim 1, wherein, The minimum value of the video-side signal-to-noise ratio and the audio-side signal-to-noise ratio is taken, and if the value is less than a first threshold value, it is determined that the value is less than the first threshold value, and then the artificial review queue is entered, wherein the first threshold value is a self-defined parameter.
6. The multi-role collaborative customer service agent method of claim 1, wherein, The video phase trajectory and the audio phase trajectory are combined to construct a linear equation, and the proportional parameter and the offset parameter of the linear equation are obtained by least square closed-form solution, and the calculation formula of the linear equation is as follows: where 1≤i≤L, L represents the total number of elements of the video phase trajectory or the audio phase trajectory, and respectively represent the i-th element value of the video phase trajectory and the audio phase trajectory, a represents a proportional parameter, and b represents an offset parameter.
7. The multi-role collaborative customer service agent method of claim 6, wherein, The residual sequence is calculated according to the video phase trajectory, the audio phase trajectory, the proportional parameter and the offset parameter, and the global consistency score is calculated according to the residual sequence, and if the global consistency score is greater than or equal to a second threshold value, it is determined that the global consistency score is greater than or equal to the second threshold value, and then the global consistency is output, otherwise the global inconsistency is output, wherein the second threshold value is a self-defined parameter; The value of the i-th element of the residual sequence The formula for calculating the value of the i-th element of the residual sequence is as follows: where 1≤i≤L, L represents the total number of elements of the video phase trajectory or the audio phase trajectory, and respectively represent the i-th element value of the video phase trajectory and the audio phase trajectory, a represents a proportional parameter, and b represents an offset parameter; Global consistency score The formula for calculating the global consistency score is as follows: wherein denotes the sample variance of the residual sequence, denotes the sample variance of the video phase trajectory. 8.The multi-role collaborative customer service agent method of claim 7, wherein, The number of sequence units of the feature sequence is the same as the total number of elements of the video phase trajectory and the audio phase trajectory, the sequence units of the feature sequence are represented by feature vectors, and the feature vectors are composed of the element value of the video phase trajectory, the element value of the audio phase trajectory, the circular domain phase difference, the video phase increment, the audio phase increment, the frequency lock error, the video phase encoding and the audio phase encoding; Phase difference of the circular domain corresponding to the feature vector of the i-th sequence unit The calculation formula is as follows: wherein denotes the value of the i-th element of the residual sequence, denotes the floor function; The video phase increment and the audio phase increment respectively represent the difference between the current element value and the previous element value of the video phase trajectory and the audio phase trajectory; The frequency lock error is equal to the video phase increment minus the product of the proportional parameter and the audio phase increment; Video phase encoding of the i-th sequence unit corresponding to the feature vector The calculation formula is as follows: wherein represents a video side signal-to-noise ratio, represents the i-th element value of a video phase trajectory, represents a video phase adjustment parameter, being a user-defined parameter; audio phase encoding of the i-th sequence unit corresponds to the feature vector The calculation formula is as follows: wherein represents the audio side signal-to-noise ratio, represents the i-th element value of the audio phase trajectory, represents the audio phase adjustment parameter, is a user-defined parameter. 9.The multi-role collaborative customer service agent method of claim 8, wherein, The phase perception model is constructed based on a gated neural network model, and the product of the weight parameter and the concentration is added to the calculation formula of the reset gate and the update gate; The calculation formula of the concentration F is as follows: wherein denotes a frequency lock error, and denotes a first weight parameter and a second weight parameter, respectively; The sample labels of the training samples used to train the phase perception model are obtained by manual annotation.
10. The multi-role collaborative customer service agent method of claim 7, wherein, The size of the sliding window is a self-defined parameter, the calculation formula of the local consistency score and the global consistency score is the same, the third threshold value is a self-defined parameter, the element value change rate of the residual sequence corresponding to the sliding window is calculated, and the frame index of the video frame corresponding to the maximum change rate is selected as the breakpoint index.
Citation Information
Patent Citations
Abnormal equipment identification method and system based on artificial intelligence and equipment operation sound
CN117672255A
Digital currency transaction system and device based on iris and block chain
CN120163582A
Audio and video stream multi-mode real-time anomaly detection method and system based on intelligent feature tracking
CN120259688A
Automatic noise monitoring system and method based on multi-sensor data fusion
CN120493030A
Conference window conversation system based on voice sensor
CN120636432A