Audio and video intelligent analysis processing method, system, device and equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SKYWORTH DISPLAY TECH CO LTD
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]本申请提供了一种音视频智能分析处理方法、系统、装置、设备及存储介质,以解决现有技术中的音视频智能分析处理方法无法满足电视对稳定、低延迟、便捷化的实时AI理解需求的问题
[0017] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application connects a smart terminal device to an external AI computing device via a USB cable. The smart terminal device collects target data to be analyzed and encoded. The encoded target data is then sent to the external AI computing device via the RNDIS virtual network interface. The external AI computing device receives the encoded target data, decodes it, and calls an AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device. The obtained AI understanding result is then sent to the smart terminal device via the RNDIS virtual network interface. The smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data. Furthermore, in this application embodiment, the smart terminal device sends the target data to be analyzed and understood to the external AI computing device via the RNDIS virtual network interface, allowing the AI computing device to perform the analysis and understanding of the target data. This method avoids the problems of inconvenient audio and video stream acquisition, high transmission latency, and complex deployment inherent in mainstream streaming solutions in the prior art, achieving stable, low-latency, and convenient real-time AI understanding of television content.
Smart Images

Figure CN122513611A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of edge computing and smart TV technology, and in particular to an audio and video intelligent analysis and processing method, system, device, equipment and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, AI technology is increasingly being used in home settings, leading to a continuous increase in the demand for AI-powered audio and video content understanding in smart TVs, especially in scenarios such as intelligent video content recognition, OCR recognition and translation of subtitles. These AI audio and video understanding tasks rely on strong computing power, but mainstream Android TV devices are constrained by cost and power consumption, making it difficult to integrate high-performance AI processors.
[0003] Therefore, in the existing technology, the AI computing power is externalized to a dedicated computing box by using an HDMI capture card to obtain video signals from the TV's HDMI interface, or by pushing the TV's audio and video streams to the computing box via a WiFi wireless network. The dedicated computing box then receives the encoded audio and video streams from the TV side and performs analysis and processing.
[0004] The aforementioned existing technical solutions all have significant drawbacks in practical applications: HDMI acquisition requires additional external hardware, resulting in high costs and limitations imposed by the HDCP copyright protection protocol, making it impossible to acquire protected content normally; WiFi and LAN network streaming methods suffer from high transmission latency, unstable bandwidth, excessive network resource consumption, and cumbersome configuration. Therefore, the mainstream solutions in the existing technologies generally suffer from inconvenient audio and video stream acquisition, high transmission latency, and complex deployment, failing to meet the television's requirements for stable, low-latency, and convenient real-time AI understanding. Summary of the Invention
[0005] This application provides an audio and video intelligent analysis and processing method, system, device, equipment, and storage medium to solve the problem that existing audio and video intelligent analysis and processing methods cannot meet the television's requirements for stable, low-latency, and convenient real-time AI understanding.
[0006] In a first aspect, this application provides an audio and video intelligent analysis and processing method, the method being applied to a smart terminal device, the smart terminal device being connected to an external AI computing device via a USB cable, the method comprising: Collect target data to be analyzed and processed by AI, and encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface, which is specified by the Remote Network Driver Interface (RNDIS). The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The smart terminal device receives the AI understanding results and performs corresponding processing operations based on the result data.
[0007] In one possible implementation of this application, the target data includes: Currently playing audio and video data; The AI understanding results include: the recognition results of a specified object in the video content and / or the recognition results of a specified object in the audio content; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
[0008] In one possible implementation of this application, the AI understanding result includes: the current audio and video content related recommendation result; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the current audio and video content association recommendation results and overlays the association recommendation results onto the currently playing video frame.
[0009] In one possible implementation of this application, the AI understanding result includes: subtitle OCR and real-time translation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the subtitle OCR and real-time translation results, and overlays the translated text onto the current video frame.
[0010] In one possible implementation of this application, the AI understanding result includes: non-compliant content detection result; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the non-compliant content detection result carrying playback control instructions. If the current video frame contains non-compliant content, the playback control instructions are executed to perform operations such as skipping the screen, blurring the screen, or blocking specified content on the currently playing video frame.
[0011] In one possible implementation of this application, the step of collecting target data to be subjected to AI analysis and processing, and encoding the target data, includes: Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; The AI understanding results include: user intent data; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the user intent data, performs corresponding operations based on the user intent data, and then responds to the user interaction operation.
[0012] In one possible implementation of this application, the step of collecting target data to be subjected to AI analysis and processing, and encoding the target data, includes: Collect target data to be analyzed and processed by AI, dynamically adjust the bitrate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
[0013] Secondly, embodiments of this application provide an audio and video intelligent analysis and processing method system, the system including an intelligent terminal device and an external AI computing power device, the intelligent terminal device being connected to the external AI computing power device via a USB cable; The intelligent terminal device is used to collect target data to be analyzed and processed by AI, and to encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface. The AI computing power device is used to receive the encoded target data, decode it, call the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and send the obtained AI understanding results to the smart terminal device through the RNDIS virtual network interface. The intelligent terminal device is also used to receive the AI understanding results and perform corresponding operations based on the result data.
[0014] Thirdly, this application provides an audio and video intelligent analysis and processing device, which is applied to a smart terminal device. The smart terminal device is connected to an external AI computing device via a USB cable. The device includes: The acquisition and encoding module is used to acquire target data to be analyzed and processed by AI, and to encode the target data. The sending module is used to send the encoded target data to an external AI computing power device through the RNDIS virtual network interface. The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The AI understanding result processing module is used by the smart terminal device to receive the AI understanding result and perform corresponding processing operations based on the result data.
[0015] Fourthly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute an audio and video intelligent analysis and processing method program stored in the memory to implement the audio and video intelligent analysis and processing method described in any one of the first aspects.
[0016] Fifthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the audio and video intelligent analysis and processing method described in any one aspect.
[0017] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application connects a smart terminal device to an external AI computing device via a USB cable. The smart terminal device collects target data to be analyzed and encoded. The encoded target data is then sent to the external AI computing device via the RNDIS virtual network interface. The external AI computing device receives the encoded target data, decodes it, and calls an AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device. The obtained AI understanding result is then sent to the smart terminal device via the RNDIS virtual network interface. The smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data. Furthermore, in this application embodiment, the smart terminal device sends the target data to be analyzed and understood to the external AI computing device via the RNDIS virtual network interface, allowing the AI computing device to perform the analysis and understanding of the target data. This method avoids the problems of inconvenient audio and video stream acquisition, high transmission latency, and complex deployment inherent in mainstream streaming solutions in the prior art, achieving stable, low-latency, and convenient real-time AI understanding of television content. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0021] Figure 1 A schematic flowchart of an embodiment of an audio and video intelligent analysis and processing method provided in this application; Figure 2 A schematic flowchart of an embodiment of another audio and video intelligent analysis and processing method provided in this application; Figure 3 A schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application; Figure 4 A schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application; Figure 5 A schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application; Figure 6 A schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application; Figure 7 A schematic diagram of the structure of an audio and video intelligent analysis and processing system provided in an embodiment of this application; Figure 8 This application provides a schematic diagram of the data flow of an audio and video intelligent analysis and processing system as an embodiment of the present application. Figure 9 A schematic diagram of an embodiment of an audio and video intelligent analysis and processing device provided in this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0024] Due to cost and power consumption limitations, mainstream Android TVs in the current technology cannot integrate high-performance AI processors. Therefore, AI computing power is externalized to dedicated computing boxes. Existing external computing boxes generally suffer from high costs, restrictions imposed by the HDCP copyright protection protocol, high transmission latency, unstable bandwidth, excessive network resource consumption, and cumbersome configuration when transmitting data. Based on this, this application provides an audio and video intelligent analysis and processing method, system, device, and electronic device.
[0025] The present application will now be described in detail through specific embodiments.
[0026] Figure 1 A schematic flowchart of an embodiment of an audio and video intelligent analysis and processing method provided in this application; refer to Figure 1 As shown in the embodiment of this application, an audio and video intelligent analysis and processing method is applied to a smart terminal device. The smart terminal device is connected to an external AI computing device via a USB cable. The method includes the following steps S10-S30: S10. The intelligent terminal device collects the target data to be analyzed and processed by AI, and encodes the target data.
[0027] S20. The encoded target data is sent to an external AI computing device through the RNDIS virtual network interface; so that the AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface.
[0028] In this embodiment, a USB virtual Ethernet interface is constructed based on the Remote Network Driver Interface Specification (RNDIS) to realize point-to-point network communication between the smart terminal device and the external edge AI computing box.
[0029] S30. The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data.
[0030] In this embodiment, the aforementioned smart terminal device may be an Android TV, a smartphone running the Android system, a tablet computer, an in-vehicle terminal, a smart projector, or other devices, and is not limited thereto.
[0031] Taking the aforementioned smart terminal device as an Android TV device as an example, the Android TV device establishes a connection with the external AI computing power device through a USB cable. The external AI computing power device, acting as the USB device end, reports the RNDIS device type to the Android TV device. Based on the USB device enumeration information, the Android TV device automatically loads the native RNDIS driver, creates an RNDIS virtual Ethernet network card, and establishes a point-to-point virtual network connection with the external AI computing power device.
[0032] During the target data encoding and transmission process, the Android TV device encapsulates the encoded target data into Ethernet data frames and sends them to the RNDIS virtual Ethernet network card through the created Socket interface. The system layer converts the Ethernet data frames into USB transmission messages and sends them to the external AI computing power device through the USB physical channel, thereby realizing the target data transmission based on the RNDIS virtual network.
[0033] Furthermore, in this embodiment, the smart terminal device and the external AI computing power device are connected via a single USB cable, eliminating the need for an HDMI capture card or network configuration. The smart terminal device natively supports RNDIS, enabling plug-and-play functionality and providing powerful edge AI understanding capabilities. Moreover, through low-latency transmission, real-time AI content understanding is achieved. Additionally, in this embodiment, the audio and video of the smart terminal device are processed locally, without uploading to the cloud, thus offering better privacy protection. Furthermore, in this mode, the external AI computing power box can independently upgrade its AI models and capabilities, supporting flexible expansion.
[0034] Figure 2 This is a schematic flowchart of another audio and video intelligent analysis and processing method provided in this application. The audio and video intelligent analysis and processing method provided in this embodiment is applied to a specified object recognition scenario. In this scenario, the target data includes: currently playing audio and video data, and the corresponding AI understanding results include: specified object recognition results of video content and / or specified object recognition results of audio content.
[0035] Reference Figure 2 As shown, in this embodiment, step S30 above, where the smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, includes the following step S201: S201. The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
[0036] The external AI computing power device is equipped with a related information database. After the specified object is identified, the related information is synchronously obtained and sent to the aforementioned smart terminal device. After obtaining the related information, the smart terminal device synchronously displays it in the currently playing video frame.
[0037] For example, the aforementioned specified object recognition includes: face recognition, object recognition, scene recognition, brand logo recognition, etc., thereby enabling smart terminal devices to: identify the actor's identity and display their information and a list of works; identify objects in the scene and provide object information display; identify the shooting location and display travel information; and identify brand logos to display product information, etc.
[0038] Figure 3 This is a schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application. The audio and video intelligent analysis and processing method provided in this embodiment is applied to a scenario of performing AI understanding of audio and video and making content recommendations. The AI understanding result includes: the current audio and video content related recommendation result.
[0039] Reference Figure 3 As shown, in step S30 above, the intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including the following step S301: S301. The smart terminal device receives the current audio and video content association recommendation result and overlays the association recommendation result onto the currently playing video frame.
[0040] In this embodiment, during the playback of video content on the smart terminal device, the currently playing video frame is encoded and sent to the AI computing power box in real time via the RNDIS virtual network interface. The AI computing power box decodes the video stream and analyzes the type, theme, scene, actors, and object information of the video content through the AI understanding engine. Based on the analysis results, it queries the local or cloud-based related content database, matches and generates a recommendation list containing similar films and television shows, works by the same actors, related products, or related news information. The recommendation list is then transmitted back to the smart terminal device via the RNDIS virtual network interface. The smart terminal device displays the recommendation list in the form of recommendation cards on the playback interface for users to quickly view and select, realizing intelligent related recommendations based on content understanding.
[0041] Figure 4 This is a schematic flowchart of another audio and video intelligent analysis and processing method provided in this application. The audio and video intelligent analysis and processing method provided in this embodiment is applied to a scenario of subtitle recognition and translation of currently playing audio and video data. The AI understanding results in the corresponding current scenario include: subtitle OCR and real-time translation results.
[0042] Reference Figure 4 As shown, in this embodiment, step S30 above, where the smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, includes the following step S401: S401. The smart terminal device receives the subtitle OCR and real-time translation results, and displays the translated text overlaid on the current video frame.
[0043] When a smart terminal device plays video content with foreign language subtitles, the video stream is sent to the AI computing box via the RNDIS virtual network interface. The AI computing box decodes the received video stream, detects and locates the subtitle area in the video frame, extracts the subtitle text using optical character recognition (OCR), and then translates the recognized foreign language subtitles into Chinese using a machine translation model. The AI computing box then sends the translation result, which includes the original text, the translated text, and the timestamp information, back to the smart terminal device via the RNDIS virtual network interface. The smart terminal device synchronizes according to the timestamp and overlays the translated Chinese subtitles onto the video frame, achieving real-time subtitle translation and display.
[0044] Figure 5 This is a schematic flowchart of another embodiment of the audio and video intelligent analysis and processing method provided in this application. The audio and video intelligent analysis and processing method provided in this embodiment is applied to the scenario of detecting non-compliant content in the played audio and video content. The AI understanding result obtained at this time includes: non-compliant content detection result.
[0045] Non-compliant content can refer to inappropriate content such as violence and terrorism, or content unsuitable for teenagers and children.
[0046] Reference Figure 5 As shown, in step S30 above, the smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including the following step S501: S501, the smart terminal device receives the non-compliant content detection result carrying the playback control instruction. If the current video frame contains non-compliant content, the playback control instruction is executed to perform screen skipping, screen blurring, or specified content blocking operations on the currently playing video frame.
[0047] In this embodiment, the AI computing power box uses an AI understanding engine to detect and analyze the video content and audio information, enabling the identification of violent scenes, terrifying content, inappropriate language, and age rating. When the computing power box detects that the current content is inappropriate or illegal, it sends control commands to the smart terminal device through the RNDIS virtual network interface, causing the TV to skip the corresponding segment, blur the image, or mute it. Alternatively, the computing power box can directly output a black screen protection signal to achieve real-time interception and viewing protection of inappropriate content, ensuring the content safety of the viewer.
[0048] Figure 6 This is a schematic flowchart of another audio and video intelligent analysis and processing method provided in this application. In this application embodiment, the intelligent terminal device collects microphone or camera data, encodes it, and transmits it to an external AI computing device through the RNDIS virtual network interface for AI understanding. The corresponding AI understanding results include: user intent data.
[0049] Reference Figure 6 As shown, step S10 above, which involves collecting target data to be analyzed and processed by AI, and encoding the target data, includes the following step S101: S101. Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; In this embodiment, step S30 above, where the smart terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, includes the following step S601: S601. The smart terminal device receives the user intent data, performs a corresponding operation based on the user intent data, and then responds to the user interaction operation.
[0050] For example, when a user needs to control a smart terminal device by voice, the user's voice data is collected through the smart terminal device's built-in microphone or an external USB microphone; the smart terminal device performs audio encoding processing on the collected voice data, and after encoding, it sends it to the AI computing power box in real time through the RNDIS virtual network interface; the AI computing power box decodes the audio data, for example, converting the speech into text through a local ASR automatic speech recognition model, and then parsing the user's control intention through an NLU natural language understanding model; the AI computing power box generates corresponding control commands based on the parsing results, and sends the control commands back to the smart terminal device through the RNDIS virtual network interface.
[0051] Smart terminal devices execute corresponding operations based on received control commands, including changing channels, searching for content, playing, pausing, and adjusting volume.
[0052] This embodiment can complete the entire voice interaction process offline without connecting to the external network, and features fast response speed, local processing of user voice data, and strong privacy protection.
[0053] Of course, this embodiment can also enable users to interact with smart terminal devices using gesture operations. In this scenario, the smart terminal device can collect image data from its built-in camera and send it to the AI computing box for recognition.
[0054] In the above embodiments of this application, the smart terminal device first sends the AI task type to the AI computing power box at the same time as or before sending the encoded data. This type of AI task type represents the task that the smart terminal device wants the AI computing power box to perform, such as face recognition, subtitle OCR recognition, children's content detection, speech recognition, etc. After the AI computing power box determines the AI task type, it loads the corresponding AI model, performs understanding and analysis on the target data according to the specified task type, and returns the results to the smart terminal device.
[0055] In one possible embodiment of this application, step S10 above, which involves collecting the target data to be processed by AI analysis and encoding the target data, includes the following step A10: Step A10: Collect the target data to be analyzed and processed by AI, dynamically adjust the bit rate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
[0056] In this embodiment, considering USB transmission latency and AI parsing latency, when the smart terminal device plays specific video content, such as when the screen change rate is lower than a specified value or higher than a specified value, the currently playing video data is collected as the target data for AI analysis and processing. The smart terminal device determines the content type based on the current video playback scenario fed back by the AI computing power box.
[0057] For example, in fast-moving, highly dynamic scenes, such as action movies or sports events, the resolution and frame rate of the target data should be increased to ensure the accuracy of AI recognition. For static images and slow-changing scenes, such as dialogues, subtitles, image displays, and still shots, reduce the bitrate, resolution, or frame rate of the target data to reduce data transmission volume and computing power consumption.
[0058] The intelligent terminal device encodes the adjusted target data and sends it to the AI computing box via RNDIS.
[0059] By using the above-mentioned scenario-adaptive adjustment method, the transmission latency and power consumption of audio and video data are reduced while ensuring the accuracy of AI understanding, thus achieving a balance between recognition effect, processing latency and system power consumption.
[0060] In a preferred embodiment of this application, for the data transmission protocol between the aforementioned smart terminal device and the AI computing box, audio and video data and control signaling are transmitted using a unified custom frame format, including: frame type, stream ID, timestamp, sequence number, data length, and payload data. The frame type includes video I / P / B frames, audio frames, control command frames, AI result frames, recommendation decision frames, control instruction frames, etc., enabling multiplexing of audio and video streams, signaling, and results within the same channel.
[0061] For example, the above protocol frame format is shown in Table 1 below:
[0062] Table 1 The aforementioned stream ID is used to distinguish different audio and video streams (such as main screen / picture-in-picture).
[0063] For example, the definitions of the above frame types are shown in Table 2 below:
[0064] Table 2 The aforementioned data is encapsulated in a layered manner from the application layer to the physical layer, adapting to the RNDIS virtual network transmission. Specifically, the application layer encapsulates custom protocol frames, carrying audio / video data, AI results, or control commands; the transport layer encapsulates the custom protocol frames using IP / UDP, adapting to the RNDIS virtual network's transmission protocol; the RNDIS layer encapsulates IP / UDP packets into RNDIS messages, transmitting them through the virtual network driver interface; and the physical layer converts RNDIS messages into USB bulk transmission messages, sending them to the peer device via a USB cable.
[0065] This layered encapsulation structure enables the orderly transmission of audio and video streams, AI results, and control signaling within the same USB-RNDIS channel, while ensuring data synchronization, link keep-alive, and service differentiation, providing reliable protocol support for low-latency, full-duplex communication.
[0066] Figure 7 This is a schematic diagram of the structure of an audio and video intelligent analysis and processing system provided in an embodiment of this application; refer to Figure 7 As shown, the audio and video intelligent analysis and processing method system includes an intelligent terminal device 701 and an external AI computing power device 702. The intelligent terminal device is connected to the external AI computing power device via a USB cable.
[0067] The aforementioned intelligent terminal device 701 is used to collect target data to be analyzed and processed by AI, and to encode the target data.
[0068] The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface.
[0069] The aforementioned AI computing device 702 is used to receive the encoded target data, decode it, call the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and send the obtained AI understanding results to the aforementioned smart terminal device 701 through the RNDIS virtual network interface. The aforementioned smart terminal device 701 is also used to receive the AI understanding results and perform corresponding operations based on the result data.
[0070] This system comprises a smart terminal device and an external AI computing power box. The two establish an RNDIS virtual network link via a USB interface to achieve full-duplex data transmission. The smart terminal device acts as the audio and video encoding source, collecting and encoding audio and video data, and sending it to the computing power box via RNDIS. The AI computing power box receives and decodes the audio and video data, performs AI content understanding, and transmits the structured understanding results back to the smart terminal device via RNDIS for display and interaction.
[0071] Figure 8This application provides a schematic diagram of the data flow of an audio / video intelligent analysis and processing system as an embodiment of the present application; see reference. Figure 8 As shown, the overall system consists of two parts: a smart terminal device (audio and video encoding source) and an AI computing power box. The two communicate in full-duplex mode through a USB-RNDIS virtual network channel. This embodiment takes an Android TV as the smart terminal device as an example. The specific process is as follows: Android TV acts as the audio and video encoding source, encoding the currently playing or captured audio and video data. Video uses H.264 / H.265 encoding format, and audio uses AAC / Opus encoding format. The encoded audio and video streams are sent to the AI computing box via the uplink USB-RNDIS channel. After receiving the audio and video streams, the AI computing box decodes them and sends them to the AI understanding engine for analysis and processing.
[0072] The AI understanding engine inside the AI computing power box performs multi-dimensional content understanding on audio and video data. The AI understanding results that can be achieved include: object detection, face recognition, scene classification, text and speech recognition, content recommendation list generation, and control command generation.
[0073] The structured AI understanding results (JSON format) generated by the AI understanding engine are transmitted back to the Android TV device via the downlink USB-RNDIS channel. After receiving the AI results, the Android TV performs corresponding UI display, content recommendation, or control operations based on the result type, such as overlaying subtitles, displaying recommendation cards, or performing content masking.
[0074] The entire process of this system achieves full-duplex data transmission through a single USB cable, transmitting audio and video streams upstream and AI understanding results downstream. No additional hardware or network configuration is required. It features low latency, driverless plug-and-play functionality, and enables efficient real-time interaction between Android TV and AI computing box.
[0075] The system provided in this embodiment has the same specific functions and working steps as described in the above method embodiment, and will not be repeated here.
[0076] Figure 9 A schematic diagram of an embodiment of an audio and video intelligent analysis and processing device provided in this application; see reference. Figure 9 As shown, this embodiment provides an audio and video intelligent analysis and processing device. The device is applied to a smart terminal device, which is connected to an external AI computing device via a USB cable. The device includes: The acquisition and encoding module 901 is used to acquire target data to be analyzed and processed by AI, and to encode the target data; The sending module 902 is used to send the encoded target data to an external AI computing power device through the RNDIS virtual network interface; The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. AI understanding result processing module 903 is used for the smart terminal device to receive the AI understanding result and perform corresponding processing operations based on the result data. In one possible implementation, the target data includes: Currently playing audio and video data; The AI understanding results include: the recognition results of a specified object in the video content and / or the recognition results of a specified object in the audio content; The aforementioned AI understanding result processing module 903 is specifically used for: The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
[0077] In one possible implementation, the AI understanding result includes: the current audio and video content related recommendation result; The aforementioned AI understanding result processing module 903 is specifically used for: The smart terminal device receives the current audio and video content association recommendation results and overlays the association recommendation results onto the currently playing video frame.
[0078] In one possible implementation, the AI understanding results include: subtitle OCR and real-time translation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the subtitle OCR and real-time translation results, and overlays the translated text onto the current video frame.
[0079] In one possible implementation, the AI understanding results include: non-compliant content detection results; The aforementioned AI understanding result processing module 903 is specifically used for: The smart terminal device receives the non-compliant content detection result. If the current video frame contains non-compliant content, it performs operations such as skipping the screen, blurring the screen, or blocking specified content on the currently playing video frame.
[0080] In one possible implementation, the acquisition and encoding module 901 is specifically used for: Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; The AI understanding results include: user intent data; The aforementioned AI understanding result processing module 903 is specifically used for: The smart terminal device receives the user intent data, performs corresponding operations based on the user intent data, and then responds to the user interaction operation.
[0081] In one possible implementation, the acquisition and encoding module 901 is specifically used for: Collect target data to be analyzed and processed by AI, dynamically adjust the bitrate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
[0082] This device uses the RNDIS protocol to virtualize the USB physical link as an Ethernet network card, enabling point-to-point full-duplex communication between the TV and the computing box. The TV can actively encode and send audio and video streams, and the computing box can transmit AI understanding results back in real time, without the need for additional hardware and network configuration.
[0083] like Figure 10 As shown in the figure, this application provides a device including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, the processor 111, when executing a program stored in the memory 113, implements the audio and video intelligent analysis and processing method provided in any of the foregoing method embodiments. The method is applied to a smart terminal device, which is connected to an external AI computing device via a USB cable, and includes: Collect target data to be analyzed and processed by AI, and encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface. The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The smart terminal device receives the AI understanding results and performs corresponding processing operations based on the result data.
[0084] In one possible implementation, the target data includes: Currently playing audio and video data; The AI understanding results include: the recognition results of a specified object in the video content and / or the recognition results of a specified object in the audio content; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
[0085] In one possible implementation, the AI understanding result includes: the current audio and video content related recommendation result; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the current audio and video content association recommendation results and overlays the association recommendation results onto the currently playing video frame.
[0086] In one possible implementation, the AI understanding results include: subtitle OCR and real-time translation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the subtitle OCR and real-time translation results, and overlays the translated text onto the current video frame.
[0087] In one possible implementation, the AI understanding results include: non-compliant content detection results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the non-compliant content detection result. If the current video frame contains non-compliant content, it performs operations such as skipping the screen, blurring the screen, or blocking specified content on the currently playing video frame.
[0088] In one possible implementation, the step of collecting target data to be analyzed and processed by AI, and encoding the target data, includes: Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; The AI understanding results include: user intent data; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the user intent data, performs corresponding operations based on the user intent data, and then responds to the user interaction operation.
[0089] In one possible implementation, the step of collecting target data to be analyzed and processed by AI, and encoding the target data, includes: Collect target data to be analyzed and processed by AI, dynamically adjust the bitrate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
[0090] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the audio-video intelligent analysis and processing method provided in any of the foregoing method embodiments: Collect target data to be analyzed and processed by AI, and encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface. The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The smart terminal device receives the AI understanding results and performs corresponding processing operations based on the result data.
[0091] In one possible implementation, the target data includes: Currently playing audio and video data; The AI understanding results include: the recognition results of a specified object in the video content and / or the recognition results of a specified object in the audio content; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
[0092] In one possible implementation, the AI understanding result includes: the current audio and video content related recommendation result; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the current audio and video content association recommendation results and overlays the association recommendation results onto the currently playing video frame.
[0093] In one possible implementation, the AI understanding results include: subtitle OCR and real-time translation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the subtitle OCR and real-time translation results, and overlays the translated text onto the current video frame.
[0094] In one possible implementation, the AI understanding results include: non-compliant content detection results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the non-compliant content detection result. If the current video frame contains non-compliant content, it performs operations such as skipping the screen, blurring the screen, or blocking specified content on the currently playing video frame.
[0095] In one possible implementation, the step of collecting target data to be analyzed and processed by AI, and encoding the target data, includes: Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; The AI understanding results include: user intent data; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the user intent data, performs corresponding operations based on the user intent data, and then responds to the user interaction operation.
[0096] In one possible implementation, the step of collecting target data to be analyzed and processed by AI, and encoding the target data, includes: Collect target data to be analyzed and processed by AI, dynamically adjust the bitrate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0099] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0100] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An intelligent audio and video analysis and processing method, characterized in that, The method is applied to a smart terminal device, which is connected to an external AI computing device via a USB cable. The method includes: Collect target data to be analyzed and processed by AI, and encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface, which is specified by the Remote Network Driver Interface (RNDIS). The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The smart terminal device receives the AI understanding results and performs corresponding processing operations based on the result data.
2. The method according to claim 1, characterized in that, The target data includes: Currently playing audio and video data; The AI understanding results include: the recognition results of a specified object in the video content and / or the recognition results of a specified object in the audio content; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the specified object recognition result of the video content and / or the specified object recognition result of the audio content, and presents the specified object recognition result and associated information together in the currently playing video frame.
3. The method according to claim 2, characterized in that, The AI understanding results include: the current audio and video content related recommendation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the current audio and video content association recommendation results and overlays the association recommendation results onto the currently playing video frame.
4. The method according to claim 2, characterized in that, The AI understanding results include: subtitle OCR and real-time translation results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the subtitle OCR and real-time translation results, and overlays the translated text onto the current video frame.
5. The method according to claim 2, characterized in that, The AI understanding results include: non-compliant content detection results; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the non-compliant content detection result carrying playback control instructions. If the current video frame contains non-compliant content, the playback control instructions are executed to perform operations such as skipping the screen, blurring the screen, or blocking specified content on the currently playing video frame.
6. The method according to claim 1, characterized in that, The process of collecting target data for AI analysis and processing, and encoding the target data, includes: Collect user interaction data locally collected by the smart terminal device and encode the user interaction data; wherein, the user interaction data includes microphone data and / or camera data collected by the smart terminal device; The AI understanding results include: user intent data; The intelligent terminal device receives the AI understanding result and performs corresponding processing operations based on the result data, including: The smart terminal device receives the user intent data, performs corresponding operations based on the user intent data, and then responds to the user interaction operation.
7. The method according to claim 2, characterized in that, The process of collecting target data for AI analysis and processing, and encoding the target data, includes: Collect target data to be analyzed and processed by AI, dynamically adjust the bitrate, resolution or frame rate of the target data according to the current video playback scenario, and encode the adjusted target data.
8. A system for intelligent audio and video analysis and processing, characterized in that, The system includes a smart terminal device and an external AI computing power device, wherein the smart terminal device is connected to the external AI computing power device via a USB cable; The intelligent terminal device is used to collect target data to be analyzed and processed by AI, and to encode the target data; The encoded target data is sent to an external AI computing device via the RNDIS virtual network interface. The AI computing device is used to receive the encoded target data, decode it, call the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and send the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The intelligent terminal device is also used to receive the AI understanding results and perform corresponding operations based on the result data.
9. An audio and video intelligent analysis and processing device, characterized in that, The device is applied to a smart terminal device, which is connected to an external AI computing device via a USB cable. The device includes: The acquisition and encoding module is used to acquire target data to be analyzed and processed by AI, and to encode the target data. The sending module is used to send the encoded target data to an external AI computing power device through the RNDIS virtual network interface. The AI computing device receives the encoded target data, decodes it, calls the AI recognition model to perform content understanding analysis according to the AI task type currently specified by the smart terminal device, and sends the obtained AI understanding result to the smart terminal device through the RNDIS virtual network interface. The AI understanding result processing module is used by the smart terminal device to receive the AI understanding result and perform corresponding processing operations based on the result data.
10. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute an audio / video intelligent analysis and processing program stored in the memory to implement the audio / video intelligent analysis and processing method according to any one of claims 1-8.
11. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the audio and video intelligent analysis and processing method according to any one of claims 1-8.