Video outbound method, device, equipment, medium and product

By automatically generating pixel-level fusion of video frames and background images and using TTS voice conversion technology, combined with VoLTE technology, the problems of low automation and poor compatibility in existing video outbound calls have been solved, achieving efficient and intuitive information transmission and improved customer experience.

CN121334337APending Publication Date: 2026-01-13CHINA UNITED NETWORK COMM CO LTD JIANGXI BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511488672.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In customer service for telecommunications operators, existing video outbound calling solutions rely on manually generated videos, resulting in low automation, poor information delivery intuitiveness, unsatisfactory customer experience, and insufficient compatibility.

Method used

By automatically generating video frames and blending them with background images at the pixel level, combined with TTS text-to-speech and VoLTE technologies, the system can automatically generate and push video content, adapting to various mobile phone models.

Benefits of technology

It improved the intuitiveness of information delivery and customer satisfaction, increased content generation efficiency, enhanced push success rate, and reduced operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334337A_ABST
    Figure CN121334337A_ABST
Patent Text Reader

Abstract

The invention discloses a video outbound method, device and equipment, a medium and a product, and relates to the field of artificial intelligence, and the method comprises the steps: determining a user background picture based on an outbound mobile phone number in a video call instruction; obtaining all video frames of a preset video; the preset video corresponds to the video call instruction; for any video frame, taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion on the background picture and the video frame to obtain a new video frame corresponding to the video frame; synthesizing the new video frames corresponding to the video frames into a video stream; converting the target character into voice by using a TTS character-to-voice technology; the target character corresponds to a preset video; and synthesizing the voice and the video stream to obtain a to-be-pushed video, and pushing the to-be-pushed video to a user side by using the VoLTE technology, so that the video to be pushed can be automatically generated, and the method and the device are compatible with various mobile phone models.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a video outbound call method and device, equipment, medium and product. BACKGROUND

[0002] In the outbound call of the communication operator customer service, the information is usually transmitted through voice or marketing, but the form of information is single, and the customer cannot intuitively understand the complex content, such as package details or operation steps, resulting in low communication efficiency. The existing short message or APP push scheme has poor effect due to information asynchronization, low customer participation, etc. In the related video outbound call process, the to-be-pushed video depends on manual generation, and the automation degree is low, and the compatibility of the to-be-pushed video to the client's mobile phone is insufficient, resulting in poor information transmission intuitiveness and poor customer experience. SUMMARY

[0003] The purpose of the present application is to provide a video outbound call method, device, equipment, medium and product, which can automatically generate a to-be-pushed video and is compatible with various mobile phone models.

[0004] To achieve the above purpose, the present application provides the following solutions: in a first aspect, the present application provides a video outbound call method, comprising: determining a user background picture based on an outbound mobile phone number in a video call instruction.

[0005] Obtaining all video frames of a preset video; the preset video corresponds to the video call instruction.

[0006] For any one video frame, taking the background picture as a background and the video frame as a foreground, pixel-level fusion is performed on the background picture and the video frame to obtain a new video frame corresponding to the video frame.

[0007] Synthesizing the new video frames corresponding to the video frames into a video stream.

[0008] Converting target text into speech using TTS text-to-speech technology; the target text corresponds to the preset video.

[0009] Synthesizing the speech and the video stream to obtain a to-be-pushed video, and pushing the to-be-pushed video to a user end using VoLTE technology.

[0010] In an embodiment, before determining the user background picture based on the outbound mobile phone number in the video call instruction, it further comprises: creating a communication channel based on the video call instruction.

[0011] In an embodiment, the new video frame corresponding to the video frame is obtained by taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion between the background picture and the video frame, specifically including: adjusting the size of the background picture according to the resolution of the video frame to obtain a processed background picture corresponding to the video frame.

[0012] The new video frame corresponding to the video frame is obtained by taking the processed background picture corresponding to the video frame as a background, taking the video frame as a foreground, and performing pixel-level fusion between the processed background picture corresponding to the video frame and the video frame.

[0013] In an embodiment, all video frames of a preset video are obtained, specifically by using a media bug mechanism to obtain all video frames of the preset video.

[0014] In an embodiment, the new video frame corresponding to the video frame is obtained by taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion between the background picture and the video frame, specifically including: adjusting the size of the background picture according to the resolution of the video frame to obtain a processed background picture corresponding to the video frame.

[0015] The processed background picture corresponding to the video frame is obtained by taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion between the processed background picture corresponding to the video frame and the video frame.

[0016] The processed background picture corresponding to the video frame is obtained by taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion between the processed background picture corresponding to the video frame and the video frame.

[0017] In an embodiment, the user background picture is determined based on an outbound phone number in a video call instruction, specifically including: determining user information according to the outbound phone number in the video call instruction.

[0018] The background picture is determined according to the user information.

[0019] In a second aspect, the application provides a video outbound device, including: a media server, the media server is used for determining a user background picture based on an outbound phone number in a video call instruction; obtaining all video frames of a preset video; the preset video corresponds to the video call instruction; for any one video frame, taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion between the background picture and the video frame to obtain a new video frame corresponding to the video frame; synthesizing the new video frames corresponding to the video frames into a video stream; using a TTS text-to-speech technology to convert target text into speech; the target text corresponds to the preset video; the speech and the video stream are synthesized to obtain a to-be-pushed video, and the VoLTE technology is used to push the to-be-pushed video to a user end.

[0020] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video outbound calling method described in any of the preceding claims.

[0021] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video outbound calling method described in any of the preceding claims.

[0022] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the video outbound calling method described in any of the preceding claims.

[0023] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a video outbound calling method, apparatus, device, medium, and product. By using a background image as the background and a video frame as the foreground, the background image and the frame are fused at the pixel level to obtain a new video frame corresponding to the video frame; the new video frames corresponding to each video frame are synthesized into a video stream; TTS text-to-speech technology is used to convert the target text into speech; the speech and video stream are synthesized to obtain the video to be pushed. The video to be pushed can be automatically generated. Because VoLTE technology is used to push the video to the user terminal, it is compatible with various mobile phone models, effectively solving the problems of poor intuitiveness of information transmission and low customer experience. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a video outbound calling method provided in one embodiment of this application.

[0026] Figure 2 This is a schematic diagram of the functional modules of a video outbound calling device provided in an embodiment of this application.

[0027] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] This application provides a video outbound calling method. In one exemplary embodiment, such as... Figure 1 As shown, the video outbound calling method includes the following steps 201 to 206.

[0031] Step 201: Determine the user's background image based on the outbound mobile phone number in the video call command.

[0032] Step 202: Obtain all video frames of the preset video; the preset video corresponds to the video call command.

[0033] Step 203: For any video frame, using the background image as the background and the video frame as the foreground, perform pixel-level fusion of the background image and the video frame to obtain a new video frame corresponding to the video frame.

[0034] Step 204: Combine the new video frames corresponding to each video frame into a video stream. Specifically, FFmpeg technology is used to combine the new video frames corresponding to each video frame into a video stream.

[0035] Step 205: Use TTS (Text-to-Speech) technology to convert the target text into speech; the target text corresponds to a preset video.

[0036] Step 206: Combine the voice and the video stream to obtain the video to be pushed, and use VoLTE technology to push the video to the user terminal.

[0037] In another exemplary embodiment of this application, before determining the user's background image based on the outbound mobile phone number in the video call instruction, the method further includes: creating a communication channel based on the video call instruction. Specifically: a video call instruction is generated according to a preset task strategy (business scenario outbound call task process configuration), a call request signaling (SIP or RTP protocol) is sent to the target user terminal, a communication channel is actively established, and the subsequent video outbound call process is triggered.

[0038] In another exemplary embodiment of this application, determining the user's background image based on the outbound mobile phone number in the video call instruction specifically includes: determining user information based on the outbound mobile phone number in the video call instruction.

[0039] The background image is determined based on the user information.

[0040] In practical applications, user information (user phone number, membership level, and region code (ISO 3166 standard)) is retrieved based on the outbound mobile phone number, generating structured data: Profile={userPhone,vipLevel,region}.

[0041] The media server calls the rules engine to execute the matching algorithm: bgID represents the background image ID, while premium_bg, campaign_bg, th_bg, and default_bg are all names of images in the background image library. Based on the returned bgID, the system retrieves the corresponding user-level image from the background image library, achieving precise personalized matching and improving user conversion rates.

[0042] In another exemplary embodiment of this application, all video frames of a preset video are obtained by using a media bug mechanism.

[0043] In another exemplary embodiment of this application, using the background image as the background and the video frame as the foreground, the background image and the video frame are pixel-level fused to obtain a new video frame corresponding to the video frame. Specifically, this includes: adjusting the size of the background image according to the resolution of the video frame to obtain a processed background image corresponding to the video frame. Specifically, this involves image scaling and adaptation: scaling or cropping the background image according to the resolution of the current video frame to ensure it matches the size of the video frame.

[0044] Using the processed background image corresponding to the video frame as the background and the video frame as the foreground, the processed background image corresponding to the video frame is merged with the video frame at the pixel level (e.g., direct replacement or on-demand mixing) to obtain a new video frame corresponding to the video frame. Specifically, the function chromakey is used to merge the background image and the video frame at the pixel level.

[0045] In practical applications, TTS (Text-to-Speech) technology is used to convert target text into speech. Specifically, TTS technology converts text into a speech PCM stream, and streaming output is used to accelerate the output speed. The TTS output process first performs front-end processing (inverse text normalization) on the text, then uses a large model to perform language inference on the front-end processing result to obtain a tensor, then inputs the tensor into a streaming model to infer a high-quality Mel spectrogram, and finally uses HiFi-GAN to process the Mel spectrogram to generate an audio waveform to obtain the speech PCM stream.

[0046] In practical applications, the pixel-level fusion process between the background image and the video frame also includes intelligent frame skipping and audio protection steps. Mechanisms such as frame skipping, non-blocking locks, and processing time budgets are used to process the video frames, ensuring that image processing does not cause audio stuttering.

[0047] In practical applications, video outbound calling methods also include: initial image setting or forced refresh optimization steps.

[0048] The purpose of the initial image setup or forced refresh optimization steps is to prioritize and quickly display the image when setting the image for the first time or when a forced refresh is performed, thereby reducing latency.

[0049] In another exemplary embodiment of this application, the voice and the video stream are combined to obtain a video to be pushed, and VoLTE technology is used to push the video to the user terminal. Specifically, the voice and the video stream are combined to obtain a video to be pushed.

[0050] Encode the video to be pushed into a transmission format (H.264 or H.265).

[0051] VoLTE technology is used to push the encoded video to the user terminal via real-time transmission protocols (RTP or WebRTC).

[0052] In practical applications, after the encoded video to be pushed to the user terminal, it also includes: synchronous monitoring of end-to-end latency (target value ≤ 300ms).

[0053] In practical applications, after the user receives the video, it calls the hardware decoder to decode the video and renders the final video frame on the display device, achieving personalized display with zero computational load.

[0054] The video outbound calling method based on VoLTE proposed in this application has the following main advantages: High content generation efficiency: Compared with the traditional outbound calling system that relies on manual generation of video or image content, this application realizes automatic content generation, which improves efficiency by more than 80% and significantly reduces operating costs.

[0055] The information delivery is highly intuitive: By using TTS (Text-to-Speech) technology and simultaneously pushing automatically generated video content to customers' mobile phones, customers can more intuitively understand complex information (such as package details or operation steps), increasing customer satisfaction by 20%.

[0056] High push success rate: The success rate of pushing video or image content to customers' mobile phones is over 95%, and it is compatible with multiple devices, solving the problems of unstable push or poor device compatibility in traditional solutions.

[0057] Based on the same inventive concept, this application also provides a video outbound calling device for implementing the video outbound calling method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more video outbound calling device embodiments provided below can be found in the limitations of the video outbound calling method described above, and will not be repeated here.

[0058] In one exemplary embodiment, such as Figure 2 As shown, a video outbound calling device is provided, comprising: a media server (a computer device with a CPU, GPU, and memory), wherein the media server is used to determine a user background image based on the outbound mobile phone number in the video call command; acquire all video frames of a preset video; the preset video corresponds to the video call command; for any video frame, using the background image as the background and the video frame as the foreground, perform pixel-level fusion of the background image and the video frame to obtain a new video frame corresponding to the video frame; synthesize the new video frames corresponding to each video frame into a video stream; use TTS text-to-speech technology to convert target text into speech; the target text corresponds to the preset video; synthesize the speech and the video stream to obtain a video to be pushed, and use VoLTE technology to push the video to be pushed to the user terminal.

[0059] In practical applications, the media server is also used to generate video call instructions according to the preset task strategy (business scenario outbound call task process configuration), send call request signaling (SIP or RTP protocol) to the target user terminal, actively establish a communication channel, and trigger the subsequent video outbound call process.

[0060] In practical applications, the video outbound calling device also includes: an outbound calling control layer module, a communication protocol module, and a user terminal.

[0061] The media server connects to the outbound call control layer module via socket, which is used to allocate computing resources (CPU cores / GPU computing power), formulate background image library rules, synthesize H264 encoded video streams, and assemble user information structures.

[0062] The outbound call control layer module's input port connects to the media server, and its output port connects to the user terminal via the communication protocol module. The outbound call control layer module is used to organize outbound user data, manage and send call sessions (up to 1000 calls), and trigger video call commands. The communication protocol module is used to parse SIP or RTP signaling (call request / response), forward media streams to the user terminal, and output SIP or RTP signaling to the user terminal. The user terminal is used to parse SIP or RTP signaling (respond to incoming calls).

[0063] The workflow of the video outbound calling device provided in this application embodiment is roughly as follows: the media server sends the video call instruction to the outbound calling control layer module, and the outbound calling control layer module sends a call request to the user terminal according to the set frequency.

[0064] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores video outbound call data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a video outbound call method.

[0065] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0066] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method embodiments.

[0067] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method embodiments.

[0068] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method embodiments.

[0069] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0070] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0071] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0072] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0073] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A video outbound calling method, characterized in that, The video outbound calling method includes: Determine the user's background image based on the outbound mobile phone number in the video call command; Retrieves all video frames of a preset video; the preset video corresponds to the video call command. For any video frame, the background image is used as the background and the video frame is used as the foreground. The background image and the video frame are merged at the pixel level to obtain a new video frame corresponding to the video frame. Combine the new video frames corresponding to each video frame into a video stream; The target text is converted into speech using TTS (Text-to-Speech) technology; the target text corresponds to a preset video. The audio and video stream are combined to obtain the video to be pushed, and VoLTE technology is used to push the video to the user terminal.

2. The video outbound calling method according to claim 1, characterized in that, Before determining the user's background image based on the outbound mobile phone number in the video call command, the following steps are also included: A communication channel is created based on video call commands.

3. The video outbound calling method according to claim 1, characterized in that, Using the background image as the background and the video frame as the foreground, the background image and the video frame are pixel-level fused to obtain a new video frame corresponding to the video frame, specifically including: The size of the background image is adjusted according to the resolution of the video frame to obtain the processed background image corresponding to the video frame; Using the processed background image corresponding to the video frame as the background and the video frame as the foreground, the processed background image corresponding to the video frame is fused with the video frame at the pixel level to obtain a new video frame corresponding to the video frame.

4. The video outbound calling method according to claim 1, characterized in that, To obtain all video frames of a preset video, specifically, a media bug mechanism is used to obtain all video frames of the preset video.

5. The video outbound calling method according to claim 1, characterized in that, The audio and video stream are combined to obtain the video to be pushed, and VoLTE technology is used to push the video to the user terminal, specifically as follows: The audio and the video stream are combined to obtain the video to be pushed; Encode the video to be pushed into a transmission format; VoLTE technology is used to push the encoded video to the user terminal via a real-time transmission protocol.

6. The video outbound calling method according to claim 1, characterized in that, Determining the user's background image based on the outbound mobile phone number in the video call command, specifically including: User information is determined based on the outbound mobile phone number in the video call instruction; The background image is determined based on the user information.

7. A video outbound calling device, characterized in that, The video outbound calling device includes: a media server, which is used to determine the user's background image based on the outbound mobile phone number in the video call command; acquire all video frames of a preset video; the preset video corresponds to the video call command; for any video frame, using the background image as the background and the video frame as the foreground, perform pixel-level fusion of the background image and the video frame to obtain a new video frame corresponding to the video frame; synthesize the new video frames corresponding to each video frame into a video stream; use TTS text-to-speech technology to convert the target text into speech; the target text corresponds to the preset video; synthesize the speech and the video stream to obtain a video to be pushed, and use VoLTE technology to push the video to be pushed to the user terminal.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the video outbound calling method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the video outbound calling method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the video outbound calling method according to any one of claims 1-6.