Video processing method and system based on end-cloud collaboration, device, and storage medium

By deploying a combined model and feature extractor on the end-side and cloud-side, the problem of poor inference effect of the mid-end-side model in the prior art is solved, and more efficient and accurate video processing effects are achieved.

WO2025112753A1PCT designated stage expired Publication Date: 2025-06-05TAOBAO CHINA SOFTWARE

Patent Information

Application Number
PCT/CN2024/116723
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-09-04
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When the existing end-cloud collaborative inference solution is used to process live videos, there is a problem that the end-side model inference is poor and the accuracy is low.

Method used

Using the video processing method based on end-cloud collaboration, the video processing method is implemented by deploying a single-modal model and a multi-modal model on the end side, and the multi-modal model on the cloud side, the accurate video processing is achieved. The specific steps include: performing preliminary feature extraction and processing on the end side. If the result is not good, uploading the feature sequence to the cloud side for multimodal processing.

Benefits of technology

It improves the accuracy and efficiency of video processing, reduces the computing burden on the cloud, and enhances the real-time and accuracy of end-side reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116723_05062025_PF_FP_ABST
    Figure CN2024116723_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video processing method and system based on end-cloud collaboration, a device, and a storage medium. In the solution provided in the embodiments of the present application, two visual feature extractors are deployed on an end side, a first visual feature extractor can extract a first visual feature sequence of a video, and a second visual feature extractor can extract a second visual feature sequence of the video. A single-modal model is deployed on the end side, and the first visual feature extractor and the single-modal model work in conjunction to complete end-side inference, thereby reducing the cloud inference burden, and further improving the end-side inference precision while the first visual feature extractor can perform personalized feature extraction on an end-side device. In addition, a multi-modal model is deployed on a cloud side, and can perform, when the single-modal model cannot obtain a processing result meeting a requirement, multi-modal inference on the basis of the visual feature sequences uploaded by the two visual feature extractors, thereby ensuring successful processing of the video, and also improving the accuracy of a cloud inference result.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing method, system, device and storage medium based on end-cloud collaboration

[0001] Cross-references

[0002] This application refers to Chinese Patent Application No. 2023116460324, filed on December 1, 2023, entitled “Video processing method, system, device and storage medium based on end-cloud collaboration”, which is incorporated into this application in its entirety by reference. Technical Field

[0003] The present application relates to the field of video processing technology, and in particular to a video processing method, system, device and storage medium based on end-cloud collaboration. Background Art

[0004] With the development of artificial intelligence (AI), neural network models are widely used, for example, to understand the content of live video. Currently, models for understanding the content of live video are deployed in the cloud. The client (i.e., the host) uploads the live video to the cloud, and the cloud understands the content of the live video based on the model and returns the results to the client.

[0005] As live video volumes increase, the cloud faces significant pressure in terms of response latency, communication, computing, and storage overhead. Consequently, a solution for device-cloud collaborative inference has emerged, where some computation is performed on the device and some on the cloud, jointly completing inference (synergistic inference). This approach can alleviate the burden on the cloud to a certain extent, but it also poses challenges such as poor inference performance on the device-side model, such as low accuracy.

[0006] Summary of the Invention

[0007] Multiple aspects of the present application provide a video processing method, system, device and storage medium based on end-cloud collaboration, which is used to more accurately process the video to be processed based on the collaboration of the end-side single-modal model and the cloud-side multi-modal model.

[0008] An embodiment of the present application provides a video processing system based on end-cloud collaboration, including: a single-modal model and a multi-channel feature extractor deployed on an end-side device, and a multi-modal model deployed on a cloud-side device; the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor, which are respectively used to extract a first visual feature sequence and a second visual feature sequence of a video to be processed, and upload them to the cloud-side device; the first visual feature extractor is also used to provide the first visual feature sequence to the single-modal model; the single-modal model is used to perform target processing on the video to be processed according to the first visual feature sequence; the multi-modal model is used to perform target processing on the video to be processed according to at least the first visual feature sequence and the second visual feature sequence when the single-modal model cannot obtain a processing result that meets the requirements.

[0009] An embodiment of the present application also provides a video processing method based on end-cloud collaboration, including: obtaining a video to be processed generated by a target end-side device, wherein a single-modal model and a multi-channel feature extractor are deployed on the target end-side device, and the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor; inputting the video to be processed into the first visual feature extractor and the second visual feature extractor for visual feature extraction to obtain a first visual feature sequence and a second visual feature sequence; inputting the first visual feature sequence into the single-modal model to perform target processing on the video to be processed; when the single-modal model cannot obtain a processing result that meets the requirements, uploading the first visual feature sequence and the second visual feature sequence to the cloud-side device, and using the multi-modal model on the cloud-side device to perform target processing on the video to be processed.

[0010] An embodiment of the present application also provides a video processing method based on end-cloud collaboration, including: obtaining a live video generated by a live broadcast application during a live broadcast, the live broadcast application running on a target end-side device, and a single-modal model and a multi-channel feature extractor deployed on the target end-side device, the multi-channel feature extractor including a first visual feature extractor and a second visual feature extractor; inputting the live video into the first visual feature extractor and the second visual feature extractor for visual feature extraction to obtain a first visual feature sequence and a second visual feature sequence; inputting the first visual feature sequence into the single-modal model, and performing target detection or content understanding on the target product in the live video; when the single-modal model cannot obtain a processing result that meets the requirements, uploading the first visual feature sequence and the second visual feature sequence to the cloud-side device, and using the multi-modal model on the cloud-side device to perform target detection or content understanding on the target product in the live video.

[0011] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; wherein the memory is used to: store one or more computer instructions; the processor is used to execute the one or more computer instructions, so as to: execute the steps in the video processing method based on end-cloud collaboration.

[0012] An embodiment of the present application also provides a computer-readable storage medium, which, when the computer program is executed by a processor, enables the processor to implement the steps in the video processing method based on end-cloud collaboration.

[0013] In an embodiment of the present application, two visual feature extractors are deployed on the terminal side, the first visual feature extractor can extract a first visual feature sequence of the video, and the second visual feature extractor can extract a second visual feature sequence of the video; a single-modal model is deployed on the terminal side, and the first visual feature extractor and the single-modal model cooperate with each other to complete terminal-side reasoning, reducing the cloud-side reasoning burden. When the first visual feature extractor can perform personalized feature extraction for the terminal-side device, the terminal-side reasoning accuracy can be further improved; in addition, a multimodal model is deployed on the cloud side. When the single-modal model cannot obtain processing results that meet the requirements, multimodal reasoning can be performed based on the visual feature sequences uploaded by the two visual feature extractors. This can not only ensure that the video is successfully processed, but also improve the accuracy of the cloud-side reasoning results. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0015] FIG1a is a schematic structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0016] FIG1b is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0017] FIG1c is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0018] FIG1d is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0019] FIG1e is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0020] FIG1f is a schematic diagram of a marker embedding provided by an exemplary embodiment of the present application;

[0021] FIG1g is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0022] FIG1h is another structural diagram of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0023] FIG1i is a schematic structural diagram of a text encoder provided by an exemplary embodiment of the present application;

[0024] FIG2 is a flow chart of a video processing method based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0025] FIG3a is another flowchart of a video processing method based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0026] FIG3 b is an architecture diagram of a video processing method based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0027] FIG3 c is a schematic diagram of sample construction provided by an exemplary embodiment of the present application;

[0028] FIG3 d is another architecture diagram of a video processing method based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0029] Figures 3e and 3f respectively show the recall rate performance of the device-side and cloud-side of the device-cloud collaboration solution provided in an embodiment of the present application within one day;

[0030] Figures 3g and 3h respectively show the memory and CPU consumption of the model inference process during a live broadcast and the model training process after the live broadcast provided by the end-cloud collaboration solution in an embodiment of the present application;

[0031] FIG4 is a schematic structural diagram of a video processing device based on end-cloud collaboration provided by an exemplary embodiment of the present application;

[0032] FIG5 is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0033] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0035] In the existing technology, the end-cloud collaborative reasoning solution can reduce the burden on the cloud to a certain extent, but there are also problems such as poor end-side model reasoning effect, such as low accuracy. For example, as shown in the schematic diagram of the end-cloud collaborative system in Figure 1a, the main idea of ​​the end-cloud collaborative system is to perform part of the calculation on the end side and part of the calculation on the cloud side, and the end and cloud jointly complete the reasoning. Specifically, single-modal models can be set on the end-side device and the cloud-side device respectively. The reasoning module in the single-modal model 1 on the end side can perform reasoning calculations on the processed data. If the judgment condition of the early exit discriminator 1 based on the early exit mechanism is met (for example, the confidence of the reasoning result is higher than the threshold), the reasoning result 1 can be output. If the judgment condition is not met, the reasoning module on the cloud-side device can continue to be reasoned. If the judgment condition of the early exit discriminator 2 based on the early exit mechanism is met, the reasoning result 2 can be output, thereby completing the reasoning based on the end side and the cloud side.

[0036] However, this existing technology has many defects. For example, both the terminal side and the cloud side are set up with single-modal models, which cannot accurately process video data that contains multiple modal data (such as text information, audio information, and image information). In addition, this existing technology often simply splits the computing task of the same model into two tasks, one executed on the cloud side and the other on the terminal side, thereby redistributing the computing load originally belonging to the terminal side device and using the computing power of the cloud side device to reduce the computing load of the terminal side device. In other words, the processing efficiency of this solution for the data to be processed is not much improved compared with the terminal side reasoning solution or the cloud side reasoning solution, and the accuracy of reasoning is still poor.

[0037] In response to the above technical problems, in some embodiments of the present application, a new end-cloud collaborative architecture is provided, in which a first visual feature extractor is added to the end side, and a multimodal model is deployed on the cloud side. The first visual feature extractor and the unimodal model cooperate with each other to complete end-side reasoning and reduce the cloud-side reasoning burden. In addition, since the first visual feature extractor has the ability to perform personalized feature extraction for end-side devices, the end-side reasoning accuracy can be improved; further, in the event of end-side reasoning failure, multimodal reasoning can be performed through the multimodal model on the cloud side, which is conducive to improving the accuracy of the cloud-side reasoning results while ensuring that the video can be successfully processed.

[0038] The technical solutions provided by each embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0039] FIG1b is a schematic diagram of the structure of a video processing system based on end-cloud collaboration provided by an exemplary embodiment of the present application. As shown in FIG1b , the video processing system includes an end-side device 10 and a cloud-side device 20. It should be noted that the number of end-side devices may be one or more, and the interaction process between the cloud-side device 20 and each end-side device 10 is similar. Therefore, for ease of description, this embodiment will be described using an example in which the video processing system includes one end-side device 10 and a cloud-side device 20, but does not limit the number of end-side devices 10.

[0040] The client-side device 10 includes, but is not limited to, smartphones, smartwatches, tablets, laptops, and smart wearable devices. The cloud-side device 20 can be implemented as a server, including conventional servers, cloud servers, cloud hosts, virtual centers, and other servers, though this embodiment does not limit this. The server device primarily comprises a processor, hard drive, memory, and system bus, similar to a common computer architecture, and will not be further described.

[0041] As shown in Figure 1b, a multi-channel feature extractor 101 and a single-modal model 102 can be deployed on the end-side device 10. Among them, the multi-channel feature extractor 101 can perform multi-modal feature extraction on the video to be processed to obtain feature sequences of multiple modalities of the video to be processed. The multi-channel feature extractor 101 may include: a first visual feature extractor 1011 and a second visual feature extractor 1012. The first visual feature extractor 1011 and the second visual feature extractor 1012 can respectively extract a first visual feature sequence and a second visual feature sequence of the video to be processed. Relatively speaking, the first visual feature extractor 1011 can be considered as a personalized visual feature extractor adapted to the end-side device 10, and the first visual feature sequence extracted by it can be considered as a personalized visual feature sequence; the second visual feature extractor 1012 is a universal visual feature extractor applicable to each different end-side device 10, and the second visual feature sequence extracted by it can be considered as a universal visual feature sequence.

[0042] In this embodiment, the first visual feature extractor 1011 can provide the first visual feature sequence to the unimodal model 102. Correspondingly, the unimodal model 102 can perform target processing on the video to be processed based on the first visual feature sequence. The target processing can be any processing method for the video, such as video classification, target detection, content understanding, video clip label generation, product identification in the video, etc., which is not limited in this embodiment. In this way, a unimodal model is deployed on the end-side device. Compared with the multimodal model, the unimodal model requires fewer computing resources and storage resources and is more suitable for the end-side device. On the one hand, with the cooperation of the unimodal model and the first visual feature extractor unimodal model that can perform personalized feature extraction for the end-side device, video processing can be completed on the end-side, and the processing results can be directly used by the end-side device, which is more real-time and can reduce the inference burden of the cloud-side device.

[0043] In this embodiment, a multimodal model 201 can be deployed on the cloud-side device 20. The multimodal model 201 can perform target processing on the video to be processed based on at least the first visual feature sequence and the second visual feature sequence when the single-modal model 102 cannot obtain the required processing results. Among them, the single-modal model 102 can obtain the corresponding processing results by performing target processing on the video to be processed based on the first visual feature sequence. For example, the single-modal model 102 can classify the video to be processed based on the first visual feature sequence to obtain a video classification result. For example, the single-modal model 102 can classify multiple video clips in the video to be processed into beautiful scenery categories, person categories, food categories, and pet categories respectively; the single-modal model 102 can also perform commodity recognition on the video to be processed based on the first visual feature sequence to obtain a commodity recognition result (commodity A1 appears in the nth frame of the video to be processed, and commodities A2 and A3 appear in the mth frame). Correspondingly, failure to obtain a processing result that meets the requirements can be understood as the confidence level of the obtained processing result being lower than the confidence level, or the single-modal model 102 failing to successfully output the processing result.

[0044] That is to say, the first visual feature extractor 1011 and the second visual feature extractor 1012 can not only respectively extract the first visual feature sequence and the second visual feature sequence of the video to be processed, but also upload the first visual feature sequence and the second visual feature sequence to the cloud-side device. In the embodiment of the present application, the timing of uploading the visual feature sequence by the visual feature extractor is not limited, and it can be the following two situations: Case 1, the visual feature extractor extracts and uploads at the same time. Case 2: The visual feature extractor first caches the extracted visual feature sequence on the end-side device 10 and uploads it to the cloud-side device 20 when the unimodal model 102 cannot obtain the required processing results. The embodiment of the present application does not limit this.

[0045] In scenario 1, the visual feature extractor can extract the visual feature sequence while uploading the extracted visual feature sequence to the cloud-side device 20. When the unimodal model 102 does not output a processing result for performing target processing on the video to be processed or does not output a processing result that meets the requirements, the end-side device 10 can send a multimodal processing notification to the cloud-side device 20, where the multimodal processing notification can be used to instruct the cloud-side device 20 to perform target processing on the video to be processed using the multimodal model 201. The cloud-side device 20 can perform target processing on the video to be processed based on the received visual feature sequence using the multimodal model 201 according to the received multimodal processing notification.

[0046] In case 2, when the unimodal model 102 does not output a processing result for the target processing of the video to be processed or does not output a processing result that meets the requirements, the first visual feature extractor 1011 and the second visual feature extractor 1012 may upload the first visual feature sequence and the second visual feature sequence to the cloud-side device for the cloud-side device to further perform target processing on the video to be processed. Specifically, when uploading the first visual feature sequence and the second visual feature sequence, the end-side device may upload the visual feature sequence cached at the historical moment and the visual feature sequence extracted at the current moment to the cloud-side device 20 together, so that the cloud-side device 20 can use the multimodal model 201 to perform target processing on the video to be processed based on the received visual feature sequence.

[0047] Among them, the target processing performed by the multimodal model 201 and the target processing performed by the unimodal model 102 are the same type of processing operations. For example, the unimodal model 102 can identify the goods in the video to be processed based on the first visual feature sequence. If it is unable to obtain the processing result that meets the requirements, for example, the accuracy of the product identification is low, then the multimodal model 201 can identify the goods in the video to be processed based on the first visual feature sequence and the second visual feature sequence. However, the multimodal model 201 and the unimodal model 102 differ in the implementation methods of performing the target processing. In this way, the multimodal model can process the video to be processed more accurately when the unimodal model cannot obtain the processing result that meets the requirements, and can reduce the computing load of the cloud-side device compared to the solution in which all video processing is performed on the cloud-side device.

[0048] Specifically, the single-modal model on the end-side device can first perform single-modal, coarse-grained processing on the video to be processed. If the processing results cannot meet the requirements, the multimodal model on the cloud-side device can further perform multimodal, fine-grained processing. Because the multimodal model has a higher computational load than the single-modal model, and the performance of the end-side device is lower than that of the cloud-side device, the multimodal model and the single-modal model are deployed on the cloud-side device and the end-side device respectively. On the one hand, the single-modal model on the end-side device can first perform single-modal, coarse-grained processing on the video to be processed, and the processing results can be directly used by the end-side device, which is more real-time. On the other hand, the multimodal model will only perform target processing on the video to be processed if the single-modal model cannot obtain the required processing results. This effectively reduces the computational load on the cloud-side device, thereby fully leveraging the respective advantages of the end-side and cloud-side devices and more accurately performing target processing on the video to be processed based on end-cloud collaboration.

[0049] In some optional embodiments, the first visual feature extractor 1011 in the aforementioned embodiment is trained based on the personalized sample data generated by the terminal device 10 where it is located. Specifically, when the user of the terminal device 10 interacts with the terminal device 10 at a historical moment, the terminal device 10 can generate personalized sample data for the user based on the user's interactive behavior. For example, when using the terminal device 10, the user can mark at least one frame of the target video with a product (i.e., mark the area where the product is located in a certain frame), and the terminal device 10 can generate personalized sample data based on the user's product marking result for the target video (i.e., the target video after the user's product marking).

[0050] Because the personalized sample data generated by the end-side device 10 is generated based on the user's historical interaction behavior, it is more in line with the user's operating habits. When the user needs to use the end-side device 10 to perform target processing on the video to be processed, the first visual feature extractor 1011 trained based on the personalized sample data can also more accurately process the video to be processed, and the processing results obtained are more in line with the user's wishes.

[0051] During the training process of the first visual feature extractor 1011, the end-side device 10 can obtain a target video that has been manually labeled with products by the user, and obtain training samples based on the target video. The format of the training samples is I = (Ia, I+, I-), where Ia is the anchor image, I+ is the positive image, and I- is the negative image. It should be noted that the product reference images corresponding to the products appearing in the target video are preset. For example, the user can preset product reference images P1 for product A1, P2 for product A2, and P3 for product A3 in the product library. Each product included in any frame of the target video can have a corresponding product reference image in the product library.

[0052] Specifically, for any frame image in at least one frame image in the target video that has been marked with a product by the user, the terminal device 10 can traverse all detection areas of the frame image; when the currently traversed detection area is closest to the product marked by the user, the detection area is used as the target detection area. If there is no detection area closest to the product marked by the user in the frame image, the frame image can be abandoned and the above operation can be performed on the next frame image. Furthermore, the terminal device can use the target detection area as Ia and the corresponding area in the product reference image corresponding to the target detection area as I+. The terminal device 10 can randomly select a detection area of ​​a product reference image that is different from the product marked by the user from all product reference images as I-; it can also select a detection area of ​​a product reference image that is closest to the aforementioned Ia from all product reference images as I-, and this embodiment does not impose any restrictions.

[0053] For any frame image in at least one frame image in the target video that has not been marked with a user product, the end-side device 10 may form multiple combinations of the detection area in the frame image and the detection area of ​​each product reference image, and sequentially calculate the matching degree of the two detection areas in each combination. Then, the end-side device 10 may use the detection area of ​​the frame image in the combination with the highest matching degree and the detection area in the corresponding reference image as Ia and I+, respectively. In addition, the end-side device 10 may randomly select a detection area of ​​a product reference image that is different from the automatically identified product as I-, or select the detection area of ​​the product reference image with the highest matching degree with Ia as I-.

[0054] Through the above method, the end-side device 10 can obtain a training sample in the format of I=(Ia, I+, I-), input the training sample into the first visual feature extractor, and obtain the output representation vector corresponding to Ia, the representation vector corresponding to I+, and the representation vector corresponding to I-. The end-side device 10 can calculate the positive example cosine distance between the representation vector corresponding to Ia and the representation vector corresponding to I+, and the negative example cosine distance between the representation vector corresponding to Ia and the representation vector corresponding to I-. Afterwards, the end-side device 10 can determine the loss of the first visual feature extractor based on the positive example cosine distance and the negative example cosine distance, and adjust the parameters of the first visual feature extractor based on the loss until the model training end condition is met, thereby obtaining the first visual feature extractor after training.

[0055] Furthermore, as video data is continuously generated, when the terminal device 10 generates new video data, new sample data can be generated based on the newly generated video data, and the first visual feature extractor 1011 can be continuously iteratively updated using the new sample data, thereby further improving the accuracy of the first visual feature extractor 1011 in extracting visual features for the terminal device 10.

[0056] In other optional embodiments, relative to the first visual feature extractor 1011, the second visual feature extractor 1012 is trained based on the unified sample data generated by multiple end-side devices 10. The unified sample data refers to the user interaction data uniformly collected from multiple end-side devices 10. In other words, the second visual feature extractor trained based on the unified sample data generated by multiple end-side devices 10 is universal and can perform unified feature extraction for each end-side device 10. However, compared to the aforementioned first visual feature extractor trained based on personalized sample data, the second visual feature extractor cannot perform personalized visual feature extraction for a certain end-side device 10. The training process of the second visual feature extractor using unified sample data is the same as the aforementioned training process of the first visual feature extractor, and will not be repeated here.

[0057] In this way, the first visual feature extractor configured on the end-side device 10 can be more adapted to the end-side device 10 than the second visual feature extractor, and can perform more personalized visual feature extraction on the video to be processed generated by the end-side device 10. In particular, in the architecture of "one cloud-side device 20 + multiple end-side devices 10", each end-side device 10 can use its personalized first visual feature extractor to extract visual features from the video to be processed, and collaborate with the unimodal model 102 to more accurately perform target processing on the video to be processed. This can give full play to the natural advantage of the end-side device being close to the user and the data source, enhance the personalized and precise reasoning capabilities of the end-side unimodal model, and relieve the pressure on the cloud-side device.

[0058] In the above or below embodiments of the present application, the implementation method of the multimodal model 201 for target processing of the video to be processed according to the first visual feature sequence and the second visual feature sequence is not limited. In some optional embodiments, the first visual feature sequence and the second visual feature sequence can be directly used as input data of the multimodal model 201, and input into the multimodal model 201 for it to perform target processing on the video to be processed. In other optional embodiments, the first visual feature sequence and the second visual feature sequence can be preprocessed to obtain target input features that conform to the multimodal model 201, and the target input features are input into the multimodal model 201 for it to perform target processing on the video to be processed according to the target input features.

[0059] In some optional embodiments, the method of preprocessing the first visual feature sequence and the second visual feature sequence to obtain the target input feature includes: using an encoder to perform an encoding operation on the first visual feature sequence and the second visual feature sequence, and then encoding the first visual feature sequence and the second visual feature sequence to obtain two encoding results, which are spliced ​​as the target input feature. In other optional embodiments, as shown in Figure 1c, the video processing system based on end-cloud collaboration may also include a prompt word generator 202 deployed on the cloud-side device 20, and the prompt word generator 202 can generate soft prompt words based on at least the first visual feature sequence and output them to the multimodal model 201. The soft prompt words (soft prompt) generated by the prompt word generator 202 can be used to prompt the multimodal model 201 with the context of the input information to guide the multimodal model 201 to output more accurately.

[0060] Further optionally, as shown in FIG1d, the prompt word generator 202 can generate a soft prompt word based on the first visual feature sequence and the second visual feature sequence. The following will describe the prompt word generation process in detail with reference to FIG1e:

[0061] As shown in Figure 1e, a first image encoder 203, a second image encoder 204 and a first splicing module 205 may also be deployed on the cloud-side device 20. The first image encoder 203 and the second image encoder 204 may be the same image encoder or different image encoders, which is not limited in this embodiment. The first image encoder 203 may encode the first visual feature sequence to obtain a first visual feature embedding vector. The second image encoder 204 may encode the second visual feature sequence to obtain a second visual feature embedding vector. The operation of encoding the visual feature sequence may be understood as a feature projection operation on the visual feature sequence, that is, transforming the visual feature sequence into a preset space to obtain a visual feature embedding vector.

[0062] The first splicing module 205 can splice the first visual feature embedding vector and the second visual feature embedding vector, and insert at least one soft prompt marker at the starting position, and insert a segmentation marker at the splicing position to obtain a third feature embedding vector. Among them, the number of soft prompt markers can be one or more, and can be customized according to design requirements, and this embodiment does not impose any restrictions. Specifically, as shown in Figure 1f, [CLS]p is a soft prompt marker that can be inserted into the head position of the first visual feature embedding vector and the second visual feature embedding vector to represent the overall meaning of the first visual feature embedding vector and the second visual feature embedding vector. The number of soft prompt markers can be determined according to the number of soft prompt words required. In Figure 1f, 2 are used as an example. [SEP]p is a segmentation marker that can be removed and inserted into the splicing position of the first visual feature embedding vector and the second visual feature embedding vector to separate the first visual feature embedding vector and the second visual feature embedding vector to represent the boundary and relationship between the two embedding vectors.

[0063] After obtaining the third feature embedding vector, the first splicing module 205 can input the third feature embedding vector into the prompt word generator 202 to generate soft prompt words. Among them, the prompt word generator 202 has been pre-trained to have the ability to generate prompt words based on the feature embedding vector. During the training of the prompt word generator 202, the cloud-side device 20 can obtain the embedding vector training sample and its corresponding prompt word label, and use the embedding vector training sample to perform multiple rounds of iterative training on the prompt word generator 202 under the supervision of the prompt word label. In the current iteration round, the cloud-side device 20 can calculate the loss function between the prompt word generated by the prompt word generator 202 and the prompt word label, and adjust the parameters of the prompt word generator 202 with the goal of converging the loss function to within a preset error range. It stops iterative training until the loss function converges to within a preset error range, and outputs the trained prompt word generator 202. In the embodiment of the application, the network architecture of the prompt word generator 202 is not limited. For example, an encoder (Encoder) in a single-layer Transformer architecture can be used, but is not limited to it.

[0064] Correspondingly, when the unimodal model 102 fails to obtain a satisfactory processing result, the multimodal model 201 may perform target processing on the video to be processed based on the first visual feature sequence and the second visual feature sequence. An optional method includes: generating initial input features based on the second visual feature sequence, embedding the generated soft prompt words into the initial input features to obtain target input features, and performing target processing on the video to be processed based on the target input features. The process of generating initial input features based on the second visual feature sequence by the multimodal model 201 will be described in detail below with reference to FIG. 1g.

[0065] As shown in FIG1g , the multi-channel feature extractor 101 may further include a text feature extractor 1013 for extracting a text feature sequence from the video to be processed and uploading it to the cloud-side device 20. The text feature extractor 1013 may perform speech recognition on the video to be processed, converting the speech content in the video into text content, and extracting a text feature sequence from the video to be processed.

[0066] Correspondingly, when the multimodal model 201 on the cloud-side device 20 generates initial input features based on the second visual feature sequence, it can generate initial input features based on the second visual feature sequence and the text feature sequence. As shown in FIG1h , the cloud-side device 20 may also be deployed with: a second image encoder 204 , a text encoder 206 , and a second splicing module 207 . The second image encoder 204 may encode the second visual feature sequence to obtain a second visual feature embedding vector. The text encoder 206 may encode the text feature sequence to obtain a text embedding vector. The structure of the text encoder 206 is shown in FIG1i . The text encoder 206 may use its internal word segmentation module to perform word segmentation on the text feature sequence, i.e., to split the text features in the text feature sequence into multiple smaller units (e.g., words or segments). The text encoder 206 may then use its internal position determination module to determine the position in the text feature sequence after word segmentation where text tags can be inserted; and use the embedding module to insert the text tags [CLS] and [SEP] at the beginning and end of the text tag sequence, respectively, to obtain a text embedding vector.

[0067] Based on this, the second concatenation module 207 may concatenate the second visual feature embedding vector and the text embedding vector to obtain the initial input features. Subsequently, the multimodal model 201 may embed the generated soft prompt words into the initial input features to obtain the target input features. In this embodiment of the present application, the network architecture of the multimodal model 201 is not limited. For example, its backbone network may adopt, but is not limited to, a multi-layer Transformer architecture.

[0068] In this way, the cloud-side device 20 can improve the accuracy of the input features of the multimodal model 201 by generating soft prompt words, thereby guiding the multimodal model 201 to perform target processing on the video to be processed more accurately based on the soft prompt words.

[0069] In addition to the video processing system based on end-cloud collaboration provided in the above embodiments, the embodiment of the present application also provides a video processing method based on end-cloud collaboration, which will be described below in conjunction with the accompanying drawings.

[0070] FIG2 is a flow chart of a video processing method based on end-cloud collaboration provided by an exemplary embodiment of the present application, which may include the steps shown in FIG2 :

[0071] Step 21: Obtain a video to be processed generated by a target end-side device. A single-modal model and a multi-channel feature extractor are deployed on the target end-side device. The multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor.

[0072] Step 22: Input the video to be processed into the first visual feature extractor and the second visual feature extractor to extract visual features, so as to obtain a first visual feature sequence and a second visual feature sequence.

[0073] Step 23: Input the first visual feature sequence into the unimodal model and perform target processing on the video to be processed.

[0074] Step 24: When the unimodal model cannot obtain a processing result that meets the requirements, the first visual feature sequence and the second visual feature sequence are uploaded to the cloud-side device, and the multimodal model on the cloud-side device is used to perform target processing on the video to be processed.

[0075] It should be noted that the execution subjects of each step of the method provided in the above embodiment can be the same device, or the method can also be executed by different devices. For example, the execution subjects of steps 21 to 24 can be the end-side device or the cloud-side device; for another example, the execution subjects of steps 21 and 22 can be the end-side device, and the execution subjects of steps 23 and 24 can be the cloud-side device; for another example, the execution subjects of steps 21 to 23 can be the end-side device, and the execution subject of step 24 can be the cloud-side device; and so on. This embodiment does not impose any restrictions. In some exemplary embodiments, the method further includes: pre-training the first visual feature extractor using personalized sample data generated by the target end-side device; and pre-training the second visual feature extractor using unified sample data generated by multiple end-side devices.

[0076] In some exemplary embodiments, a multimodal model on a cloud-side device is used to perform target processing on a video to be processed, including: inputting at least a first visual feature sequence into a prompt word generator on the cloud-side device to generate soft prompt words, and providing the generated soft prompt words to the multimodal model; generating initial input features based on a second visual feature sequence, and inputting the initial input features into the multimodal model; in the multimodal model, embedding the soft prompt words into the initial input features to obtain target input features, and performing target processing on the video to be processed based on the target input features.

[0077] In some exemplary embodiments, at least the first visual feature sequence is input into a prompt word generator on a cloud-side device to generate soft prompt words, including: inputting the first visual feature sequence and the second visual feature sequence into the prompt word generator to generate soft prompt words.

[0078] In some exemplary embodiments, the first visual feature sequence and the second visual feature sequence are input into a prompt word generator to generate soft prompt words, including: encoding the first visual feature sequence using a first image encoder to obtain a first visual feature embedding vector; encoding the second visual feature sequence using a second image encoder to obtain a second visual feature embedding vector; splicing the first visual feature embedding vector and the second visual feature embedding vector, and inserting at least one soft prompt marker at the starting position, and inserting a segmentation marker at the splicing position to obtain a third feature embedding vector; inputting the third feature embedding vector into the prompt word generator to generate soft prompt words, to obtain at least one soft prompt word corresponding to at least one soft prompt marker.

[0079] In some exemplary embodiments, the multi-channel feature extractor also includes a text feature extractor, and the method also includes: inputting the video to be processed into the text feature extractor for text feature extraction to obtain a text feature sequence, and uploading it to the cloud side device; generating initial input features based on the second visual feature sequence, including: generating initial input features based on the second visual feature sequence and the text feature sequence.

[0080] In some exemplary embodiments, initial input features are generated based on a second visual feature sequence and a text feature sequence, including: encoding the second visual feature sequence using a second image encoder on a cloud-side device to obtain a second visual feature embedding vector; encoding the text feature sequence using a text encoder on the cloud-side device to obtain a text embedding vector; and concatenating the second visual feature embedding vector and the text embedding vector to obtain the initial input features.

[0081] Based on the above embodiments, optionally, when the electronic device obtains the video to be processed generated by the target end-side device, it can obtain the live video generated by the live broadcast application on the target end-side device during the live broadcast process as the video to be processed, and the live video includes the target product, and the first visual feature sequence and the second visual feature sequence are the visual features of the target product.

[0082] It should be noted that the user can pre-set multiple candidate products before the live broadcast, and the electronic device can respond to the user's candidate product setting operation to generate product description information for the candidate products. The product description information may include at least one of the following: image, category, name and description information. The product involved in the live video generated during the live broadcast is at least one of the multiple candidate products. The electronic device can pre-extract visual features of the image of the candidate product using the first visual feature extractor and the second visual feature extractor respectively to obtain the third visual feature sequence and the fourth visual feature sequence of the candidate product, establish a correspondence between the third visual feature sequence and the product description information, and provide the correspondence to the unimodal model. For example, before the live broadcast, the user pre-sets the product description information of candidate product 1, candidate product 2 and candidate product 3. The electronic device can use the first visual feature extractor to extract visual features of the images of these three candidate products in advance, and obtain the third visual feature sequence corresponding to the product description information of candidate product 1, the third visual feature sequence corresponding to the product description information of candidate product 2, and the third visual feature sequence corresponding to the product description information of candidate product 3, and send the correspondence between these three third visual feature sequences and the product description information to the unimodal model, so that the unimodal model can perform target processing on the processed video based on the correspondence, such as target detection, category recognition or content understanding.

[0083] Correspondingly, the electronic device can input the first visual feature sequence into the unimodal model. When performing target processing on the video to be processed, the first visual feature sequence can be input into the unimodal model, and target detection or content understanding can be performed on the live video based on the first visual feature sequence and the third visual feature sequence of the candidate product.

[0084] The unimodal model traverses the third visual feature sequence of each candidate product and calculates the similarity between the currently traversed third visual feature sequence and the first visual feature sequence. The unimodal model then selects the candidate product corresponding to the third visual feature sequence whose similarity to the first visual feature sequence exceeds a similarity threshold as the target product, thereby completing object detection or content understanding.

[0085] In some optional embodiments, the electronic device may also upload the third visual feature sequence and the fourth visual feature sequence of the candidate product to the cloud-side device. Accordingly, when the electronic device uses the multimodal model on the cloud-side device to perform target processing on the video to be processed, it may be implemented based on the following steps:

[0086] Step 241: Input the first visual feature sequence and the second visual feature sequence into a multimodal model to generate multimodal data corresponding to the target product. The multimodal model may perform feature fusion on the first visual feature sequence and the second visual feature sequence to obtain the multimodal data corresponding to the target product.

[0087] In some optional feature fusion methods, the multimodal model may separately encode the first visual feature sequence and the second visual feature sequence to map the first visual feature sequence and the second visual feature sequence into the same shared feature space, thereby completing feature fusion of the first visual feature sequence and the second visual feature sequence. In other optional feature fusion methods, the multimodal model may also perform feature concatenation on the first visual feature sequence and the second visual feature sequence to complete feature fusion.

[0088] Step 242: Generate multimodal data for the candidate product based on the third and fourth visual feature sequences of the candidate product and the product description information stored on the cloud-side device. The multimodal model may perform feature fusion on the third and fourth visual feature sequences and the product description information to obtain multimodal data corresponding to the target product.

[0089] In some optional feature fusion methods, the multimodal model can separately encode the third visual feature sequence, the fourth visual feature sequence, and the product description information to map the third visual feature sequence, the fourth visual feature sequence, and the product description information into the same shared feature space, thereby completing feature fusion in the shared feature space. In other optional feature fusion methods, the multimodal model can also perform feature splicing on the third visual feature sequence, the fourth visual feature sequence, and the product description information to complete feature fusion.

[0090] Step 243: Perform target detection or content understanding on the target product based on the multimodal data of the target product and the multimodal data of the candidate products. The multimodal model may calculate the semantic feature distance between the multimodal data of the target product and the multimodal data of each candidate product, and select candidate products whose semantic feature distance from the multimodal data of the target product meets the set conditions as the products corresponding to the target product to complete target detection. The multimodal model may also use the product description information of candidate products whose semantic feature distance from the multimodal data of the target product meets the set conditions as the information understood from the target product to complete content understanding.

[0091] In this way, electronic devices can use the multimodal model on the cloud-side device to perform target processing on the video to be processed more accurately.

[0092] It is noted that the video processing solution based on end-cloud collaboration provided by the embodiments of this application can be applied to various video scenarios, such as live broadcast scenarios, short video scenarios, etc. To facilitate understanding of the technical solution of this application, the working principle of the end-cloud collaboration architecture provided by this application is described in detail below, taking the live broadcast scenario as an example.

[0093] In live streaming scenarios, in order to timely and accurately guide consumers to the product clips promoted during the live broadcast, it is necessary to have a good understanding of the live broadcast content and output high-quality and useful content understanding tags for consumers. These content understanding tags include, but are not limited to, the product name, category, brand, appearance, function, price, etc. Comparing short video content understanding with live streaming content understanding, short videos are generally less than one minute in length on average, and can be understood using offline services. However, live streaming videos average around three hours in length. Using sparse sampling for content understanding of live streaming videos would affect the real-time performance of the content understanding chain, resulting in untimely and inaccurate recommendations for consumers, which would affect conversion rates and user retention time. Therefore, real-time and full-time content understanding of live streaming videos is necessary.

[0094] Traditional solutions are generally based on the cloud service framework, that is, a model for content understanding is deployed on a cloud server. Each anchor's live broadcast device uploads the live video to the cloud server, and uses the cloud model to understand the content of the live video. Compared with terminal devices, cloud servers have more resources and stronger computing power, and to a certain extent, they can understand the content of live videos in real time and at all times.

[0095] However, with the development of live streaming technology, the number of concurrent live streamers is increasing, live streams are becoming longer, and the content is becoming richer and more diverse. This has led to several practical issues for cloud-based service frameworks, including high response latency and the real-time load and high costs of communication, computing, and storage overhead. To address these issues, strict limits can be placed on the frequency of service requests per streamer. However, this severely undermines the ability to fully understand the live content, making some content (such as products) in the live video impossible to identify and understand in a timely manner.

[0096] In response to the above-mentioned technical problems existing in the above-mentioned live broadcast scenario, this embodiment proposes a video processing architecture based on end-cloud collaboration, deploying a multimodal model on the cloud, deploying a unimodal model on the anchor end, and simultaneously deploying a general feature extractor and a personalized feature extractor suitable for each anchor end on the anchor end. The personalized feature extractor and the unimodal model cooperate with each other to first perform coarse-grained unimodal processing on the video to be processed. If the required processing result cannot be obtained, the multimodal model on the cloud will further perform fine-grained multimodal processing based on the visual feature sequences extracted by the two feature extractors. On the one hand, considering the low performance of the host's equipment, a single-modal model is used on the host side. The single-modal model can be used to perform single-modal coarse-grained processing on the video to be processed first, and the processing results can be directly used by the end-side device. With the advantage of the single-modal model being closer to the user and the data source, it can provide higher real-time response. On the other hand, considering the low accuracy of the single-modal model, a personalized feature extractor is added to the host side to provide more accurate visual features for the single-modal model, thereby improving the recognition accuracy of the single-modal model. On the other hand, with the resource and performance advantages of the cloud-side equipment, multimodal model, and the multimodal model performs target processing on the video to be processed only when the single-modal model cannot obtain the required processing results. This effectively solves the overhead problems faced by cloud-side devices in communication, computing, and storage, and reduces the computing load of cloud-side devices. When the single-modal model on the end side cannot obtain the required processing results, the multimodal model further performs multimodal processing to improve the recognition accuracy of the model. It can make full use of the respective advantages of the end-side devices and the cloud-side devices, and perform target processing on the video to be processed more accurately based on the end-cloud collaboration, which effectively solves the technical problems existing in the existing technologies in the above-mentioned live broadcast scenarios.

[0097] The above process of this solution will be further explained below in conjunction with Figures 3a, 3b, 3c and 3d using a live broadcast scenario.

[0098] As shown in FIG3a , a video processing method for a live broadcast scenario provided in an embodiment of the present application includes:

[0099] Step 31: Obtain the live video generated by the live broadcast application during the live broadcast. The live broadcast application runs on the target end-side device. A single-modal model and a multi-channel feature extractor are deployed on the target end-side device. The multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor.

[0100] Step 32: Input the live video into the first visual feature extractor and the second visual feature extractor to extract visual features, so as to obtain a first visual feature sequence and a second visual feature sequence.

[0101] Step 33: Input the first visual feature sequence into the unimodal model to perform target detection or content understanding on the target product in the live video.

[0102] Step 34: When the unimodal model cannot obtain a processing result that meets the requirements, the first visual feature sequence and the second visual feature sequence are uploaded to the cloud-side device, and the multimodal model on the cloud-side device is used to perform target detection or content understanding on the target product in the live video.

[0103] The implementation of each step in the method shown in FIG3a is described in detail below with reference to FIG3b, FIG3c and FIG3d.

[0104] Figure 3b is a schematic diagram of the end-cloud collaborative architecture provided by the embodiment of this application, deployed and implemented in a live broadcast scenario. As shown in Figure 3b, unlike traditional cloud-based service frameworks, this embodiment of the application leverages the distributed computing capabilities of numerous host terminals, taking advantage of the host terminal's natural proximity to data sources and users, allowing the host terminal and the cloud to jointly complete the task of identifying products in live broadcast videos. This not only reduces response latency, but also reduces cloud computing, communication, and storage overhead, alleviating the burden on the cloud. Specifically, as shown in Figure 3b, the architecture includes a multimodal model deployed in the cloud and a unimodal model deployed on the host side. In Figure 3b, the unimodal model on the host side is represented by matching in a unimodal product pool. In addition, two visual extractors and an automatic speech recognition module (i.e., the text feature extractor mentioned above) are also deployed on the host side. In Figure 3b, one visual extractor is a general visual extractor denoted as φ (i.e., the second visual feature extractor mentioned above), and one visual extractor is a personalized visual extractor denoted as φ1 and φ2 (i.e., the first visual feature extractor mentioned above). Among them, visual extractors φ1 and φ2 are personalized visual extractors on the two host sides. In Figure 3b, two host sides are used as an example for illustration, but this is not limited to this.

[0105] Regardless of the host, their live video is firstly processed through the automatic speech recognition module for speech recognition to obtain text features. For example, in the lower left host in Figure 3b, speech recognition obtains the text message "This is a pair of high-waisted jeans," while in the lower right host, speech recognition obtains the text message "I am carrying a beautiful envelope." Secondly, the live video is processed through an object detector for object detection. This detected live video is then fed into two visual extractors for visual feature sequence extraction. In the lower left host in Figure 3b, the target detected item is a pair of high-waisted jeans, while in the lower right host, the target detected item is an envelope. The visual extractor φ extracts a general visual feature sequence, while the visual extractors φ1 and φ2 respectively extract personalized visual feature sequences. The personalized visual feature sequences extracted by the visual extractors φ1 and φ2 are fed into the unimodal models of their respective hosts. The unimodal models match the personalized visual feature sequences against the unimodal product pool to obtain product content and list information such as the product name and category. The unimodal product pool stores the personalized visual feature sequences and corresponding product understanding for each candidate product that the host needs to promote during the current live broadcast. Furthermore, the personalized visual feature sequences for each candidate product in the unimodal product pool are pre-determined by the host's object detector, which performs object detection on images or video clips of the candidate products and then feeds them into the personalized visual extractor for visual feature extraction.

[0106] Furthermore, as shown in Figure 3b, the text information obtained by speech recognition and the visual feature sequence extracted by the two-way visual extractor are also uploaded to the cloud. When the single-modal model on the host side is unable or unable to accurately understand the content of the product, the multimodal model on the cloud side will use the text information and two-way visual feature sequence, combined with the multimodal product pool, to perform multimodal content understanding of the product. The content understanding results are further returned to the host side, and the host side can display the content understanding results on the screen of the live video. For example, the multimodal model can generate multimodal features of the product in the current live broadcast based on the text information and two-way visual feature sequence of the product, and then match the multimodal features of the product with the multimodal features of each candidate product in the multimodal product pool, and use the description information of the matching candidate product as the content understanding result of the product.

[0107] It is explained here that the multimodal product pool in the cloud stores the multimodal features of each candidate product that the anchor needs to promote during the current live broadcast and the corresponding product understanding content. The multimodal features of each candidate product include at least the general visual feature sequence, personalized visual feature sequence and text feature of the candidate product; among them, the target detector can be used in advance to perform target detection on the picture or video clip of the candidate product and then send it to the general visual extractor for visual feature extraction to obtain the general visual feature sequence of the candidate product; accordingly, the target detector can be used in advance to perform target detection on the picture or video clip of the candidate product and then send it to the personalized visual extractor for visual feature extraction to obtain the personalized visual feature sequence of the candidate product; text information can be extracted in advance from the picture or video clip or other descriptive information of the candidate product to obtain the text feature of the candidate product.

[0108] Before the live broadcast begins, users can add candidate products to be promoted in the live broadcast. The anchor can use two visual extractors to extract visual features of the candidate products. On the one hand, the personalized visual feature sequence of the candidate products is saved in the single-modal product pool. On the other hand, the two visual feature sequences are uploaded to the cloud. The cloud will store these two visual feature sequences as part of the multimodal features of the candidate products in the multimodal product pool on the cloud.

[0109] It can be seen that in the embodiment of the present application, the visual feature extraction stage is sent down to the terminal device of each anchor, and a coarse-grained product pre-recognition stage is further introduced, that is, a stage of single-modal recognition based on personalized visual features. In the pre-recognition stage, in order to reduce the amount of calculation and improve the recognition efficiency, shorter live broadcast segments are usually used. For example, live broadcast segments within 4 seconds can be used. Only a small number of live broadcast segments may not achieve high-confidence recognition. Therefore, when the single-modal model cannot obtain the required recognition results, the cloud-based multimodal model can be used to perform fine-grained multimodal recognition based on longer video features. For example, the multimodal features used for multimodal recognition in the cloud come from longer live broadcast segments, such as live broadcast segments within 12 seconds. In short, the personalized visual feature sequence used by the end-side single-modal model comes from a shorter live broadcast segment, and the two-way visual feature sequence and text information used by the cloud-side multimodal model come from a longer live broadcast segment. By introducing single-modal coarse-grained pre-recognition on the end side and retaining multi-modal fine-grained recognition on the cloud side, task collaboration and feature collaboration between the end and the cloud are achieved, thereby reducing the corresponding latency and cloud resource consumption.

[0110] Specifically, the visual extractor φ shown in Figure 3b is trained based on a variety of live videos generated by multiple anchor terminals as unified sample data, while the visual extractors φ1 and φ2 are trained based on the live videos generated by their anchor terminals as personalized sample data. Therefore, they are personalized visual extractors for each anchor.

[0111] Since the training samples used by visual extractors φ1 and φ2 are the live broadcast videos of each anchor, they have the ability to perform personalized visual feature extraction for their respective anchors. Considering the large number of anchors, a personalized visual extractor must be deployed on each anchor side. Therefore, the training of personalized visual extractors can be completed on the anchor side, which can reduce the model training load on the cloud.

[0112] The following, combined with Figure 3c, describes the sample data construction process for training visual extractors φ1 and φ2. After each live broadcast, model training can be performed based on the live video data generated during that broadcast. The sample format is I = (Ia, I+, I-), consisting of an anchor image Ia, a positive image I+, and a negative image I-. This sample data needs to be labeled, either manually or through automated recognition.

[0113] Manual labeling: The host manually labels live frames containing products. The detection area of ​​the marked live frame is then matched against the detection areas of all candidate product reference images in the live broadcast until a reference image with a detection area closest to the detection area in the marked live frame is found. The manually marked live frame serves as the anchor image Ia, and the reference image with the detection area closest to the manually marked live frame is used as the positive example image I+. For the negative example image I-, a reference image of a candidate product that is not close to the detection area in the live frame marked by the host can be randomly selected as a simple example, or a reference image of another candidate product whose detection area is closest to the anchor image Ia can be selected as a difficult example. The ratio between simple and difficult examples can be pre-set based on application requirements. If no area of ​​the frame is found that is closest to a detection area in the reference image of the marked product, the live frame is skipped. The detection area refers to the image area in the reference image or live frame that contains the product. When searching for a reference image whose detection area is closest to the detection area in the marked live frame, or when selecting reference images of other candidate products whose detection areas are closest to the anchor image Ia, the distance between the detection areas of the two images can be calculated based on the feature information of the two images, and the reference image whose detection area is closest to the detection area in the marked live frame or the reference image of other candidate products whose detection areas are closest to the anchor image Ia can be selected based on the distance. The smaller the distance between the two images, the higher the similarity between the two images. As shown in FIG3c, when the distances between the anchor image and each reference image are calculated to be 0.76, 0.88, 1.02, 1.06, 0.96, etc., respectively, the image corresponding to 0.76 is selected as the positive example image closest to the anchor image. Optionally, the above distance can be, but is not limited to, Euclidean distance.

[0114] Automatic recognition method: During the live broadcast process, a single-modal model on the host side or a multimodal model on the cloud side can be used to provide a content understanding result for a given live broadcast frame. The detection area in the given live broadcast frame can be used as the anchor image Ia, and the detection area in the reference image of the candidate product corresponding to the content understanding result can be used as the positive image I+. For the negative image I-, a detection area in the reference image of a candidate product can be randomly selected, or a detection area in the reference image of a candidate product that is automatically identified and different from the content understanding result can be selected, or a detection area in the reference image of the candidate product that is closest to the anchor image Ia can be selected. For the automatic recognition method, the inference process in the live broadcast can be reused to construct training samples, and then the model training can be performed.

[0115] In this embodiment of the present application, not only a multimodal model is deployed on the cloud, but also a prompt generation module (i.e., the prompt word generator mentioned above). The following describes the structure of the cloud-side multimodal model and prompt generation module in conjunction with Figure 3d. As shown in Figure 3d, the backbone network (referred to as the backbone) and the prompt generation module of the multimodal model are shown.

[0116] First is the backbone network of the multimodal model. In the embodiment of the present application, the architecture of the backbone network of the multimodal model is not limited. Optionally, in order to process the visual feature sequence and text information uploaded from the anchor end, as shown in FIG3d, the visual feature sequence includes a unified visual feature sequence from the visual extractor φ and a personalized visual feature sequence from the visual extractor φ1 or φ2, and the text information comes from the automatic speech recognition module of the anchor end. In FIG3d, “This is a pair of high-waisted jeans” is used as an example for illustration. In this embodiment, a single-stream backbone network of the multimodal model is used, that is, the visual feature sequence and text information are input as one input to the backbone network of the multimodal model, rather than multiple inputs.

[0117] In order to input the visual feature sequence and text information as one input to the backbone network of the multimodal model, both are preprocessed. (1) For the text feature sequence input, as shown in Figure 3d, the text information recognized from the audio information of the host in the live broadcast can be segmented using a text encoder, for example, into "this", "is"... "pants", and special tags [CLS] and [SEP] are added to the beginning and end of the text tag sequence. The text tags and their positions in the sequence are embedded together by the text encoder. The text encoder is a module before the backbone network, which is used to receive text information on the end side. It is responsible for segmenting the text information and adding text tags to obtain a text embedding vector. As shown in Figure 3d, the text embedding vector includes E [CLS] E 这 ...E 裤 E [SEP] (2) For the unified visual feature sequence, the image encoder is used to embed the unified visual feature sequence and output a unified visual embedding vector. As shown in Figure 3d, the unified visual embedding vector includes E [区域1] E [区域2] ...E [区域n] Among them, the position of the visual feature embedding is linearly mapped to the same dimension as the text embedding. The text embedding vector and the unified visual embedding vector are concatenated and input into the backbone network. As shown in Figure 3d, the concatenated feature vector includes the text embedding vector, the unified visual embedding vector and the text tag, i.e., E [CLS] E 这 ...E 裤 E [SEP] E [区域1] E [区域2] ...E[区域n] Optionally, the backbone network of the multimodal model can adopt a multi-layer Transformer architecture, and the text embedding vector and the unified visual embedding vector can be concatenated and input into the Encoder in the multi-layer Transformer to finally output a cross-modal representation of the live broadcast segment.

[0118] As shown in Figure 3d, in order to integrate the personalized visual feature sequences uploaded by each anchor end into the cloud multimodal model, the embodiment of the present application proposes a pluggable prompt generation module for automatically learning soft prompt words from the personalized visual feature sequence. The soft prompt words are further added before the text embedding vector and the unified visual embedding vector of the backbone network of the multimodal model. In addition, the unified visual feature sequence uploaded on the end side can be input into the prompt generation module together to learn the soft prompt words, which can improve the learning performance of the prompt generation module. Specifically, before the prompt generation module, two image encoders are used to project (feature transformation) the unified visual feature sequence and the personalized visual feature sequence respectively to obtain two visual embedding vectors, which are recorded as the unified visual embedding vector and the personalized visual embedding vector; the two visual embedding vectors are spliced ​​together, and the [CLS]p and [SEP]p markers are inserted between the two to obtain the feature vector E shown in Figure 3d. [CLS]p E [CLS]p E [区域1] ...E [区域n] E [SEP]p E' [区域1] ...E' [区域n] , the feature vector is input into the prompt generation model. In this embodiment, the network architecture of the prompt generation module is not limited. Optionally, a single-layer Transformer architecture can be used. The feature vector can be input into the encoder in the single-layer Transformer to learn the soft prompt word, and finally the soft prompt word E' is output. [CLS]p In the embodiment of the present application, [CLS]p is a soft prompt token, which is different from the text token [CLS]. There can be multiple soft prompt tokens, depending on the number of soft prompt words required. Among them, the learned representation of [CLS]p captures the overall meaning of the two visual feature sequences, while [SEP]p helps to understand the boundary between personalized and unified visual embedding sequences and capture the relationship between them. The final output soft prompt word E' [CLS]p is inserted into the input of the multimodal model, in Figure 3d, to output the soft prompt word E' [CLS]p Insert E [CLS] after.

[0119] In the end-cloud co-evolution framework provided in the embodiment of the present application, a unimodal model is deployed on each anchor side, which can process most live frames and upload the extracted unimodal features to the cloud. Secondly, in the embodiment of the present application, the anchor's manual labeling behavior is used on the end side to construct samples and incrementally train the unimodal model to adapt to the heterogeneous and dynamic live content of different live anchors. In addition, considering that the personalized unimodal features on the end side are inconsistently distributed in the feature space and cannot be directly integrated into the multimodal model on the cloud, the embodiment of the present application also proposes a pluggable prompt generation module to convert personalized unimodal features into prompt embeddings, and further add them to the original input of the multimodal backbone network to guide the feature fusion of specific anchors. Through this end-cloud collaboration method, not only can the cloud-side reasoning burden be reduced, and the problems of response delay and high computing, communication and storage overhead costs faced by the cloud side be solved, but the end-side reasoning accuracy can also be improved.

[0120] In order to verify the beneficial effects of the end-cloud collaboration framework provided by the embodiment of the present application in the live broadcast scenario, an online experiment was conducted. During the online experiment, the implementation scheme of the end-cloud collaboration frame provided by the embodiment of the present application is as follows: FCOS is used as the target detector on the end side, MobileNetV2 is used as the unimodal retrieval model, and the retrieval model is implemented for thousands of people and thousands of models on the anchor side; the UNITER model is used on the cloud side as the backbone of the multimodal model. Specifically, during the live broadcast of each anchor, manual or automatic product recognition will be triggered, and training samples will be constructed and saved on the local device. When the live broadcast ends, the collected samples are fine-tuned to optimize the visual feature extractor. The new feature extractor will be used in the next live broadcast and will serve as the starting point for the next round of training.

[0121] Among them, the online experimental results verified the feasibility and efficiency of the end-cloud collaborative evolution framework provided by the embodiment of the present application from multiple angles such as model accuracy, communication overhead, computing overhead and memory overhead. Specifically, (1) the end-cloud collaborative framework saved 86% of online container resources. (2) The recall rate of product recognition based on pure personalized visual features on the end side increased by 0.98%, and the recall rate of multimodal product recognition after integrating personalized visual features on the cloud increased by 14.67%; Figure 3e and Figure 3f respectively show the recall rate performance of the end side and the cloud side within one day; (3) Taking into account the resource limitations of the end side, the memory usage during the end-side model training is controlled within a reasonable range to ensure the stability of the live broadcast. Figure 3g and Figure 3h respectively show the memory and CPU consumed by the model inference process during a live broadcast and the model training process after the live broadcast. (4) Compared with the traditional solution of uploading live video to the cloud for feature extraction and voice recognition, the framework provided by the embodiment of the present application significantly reduces communication overhead by offloading voice recognition and visual feature extraction to each anchor terminal, uploading the text information, unified visual features and personalized visual features of the products in the live frame, and personalized visual features of the candidate products.

[0122] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The sequence numbers of the operations, such as 21, 22, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.

[0123] The application embodiment also provides a video processing device based on end-cloud collaboration, as shown in Figure 4, the device includes: a video acquisition unit 401, used to: acquire the video to be processed generated by the target end-side device, the target end-side device is deployed with a single-modal model and a multi-channel feature extractor, the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor; a feature extraction unit 402, used to: input the video to be processed into the first visual feature extractor and the second visual feature for visual feature extraction to obtain a first visual feature sequence and a second visual feature sequence; a single-modal processing unit 403, used to: input the first visual feature sequence into the single-modal model, and perform target processing on the video to be processed; a multi-modal processing unit 404, used to: when the single-modal model cannot obtain a processing result that meets the requirements, upload the first visual feature sequence and the second visual feature sequence to the cloud-side device, and use the multi-modal model on the cloud-side device to perform target processing on the video to be processed.

[0124] Further optionally, the feature extraction unit 402 is also used to: pre-train the first visual feature extractor using personalized sample data generated by the target end-side device; and pre-train the second visual feature extractor using unified sample data generated by multiple end-side devices.

[0125] Further optionally, when the multimodal processing unit 404 uses the multimodal model on the cloud-side device to perform target processing on the video to be processed, it is specifically used to: input at least the first visual feature sequence into the prompt word generator on the cloud-side device to generate soft prompt words, and provide the generated soft prompt words to the multimodal model; generate initial input features according to the second visual feature sequence, and input the initial input features into the multimodal model; in the multimodal model, embed the soft prompt words into the initial input features to obtain target input features, and perform target processing on the video to be processed according to the target input features.

[0126] Further optionally, when the multimodal processing unit 404 inputs at least the first visual feature sequence into the prompt word generator on the cloud-side device to generate soft prompt words, it is specifically used to: input the first visual feature sequence and the second visual feature sequence into the prompt word generator to generate soft prompt words.

[0127] Further optionally, when the multimodal processing unit 404 inputs the first visual feature sequence and the second visual feature sequence into the prompt word generator to generate soft prompt words, it is specifically used to: encode the first visual feature sequence using a first image encoder to obtain a first visual feature embedding vector; encode the second visual feature sequence using a second image encoder to obtain a second visual feature embedding vector; splice the first visual feature embedding vector and the second visual feature embedding vector, and insert at least one soft prompt marker at the starting position, and insert a segmentation marker at the splicing position to obtain a third feature embedding vector; input the third feature embedding vector into the prompt word generator to generate soft prompt words to obtain at least one soft prompt word corresponding to the at least one soft prompt marker.

[0128] Further optionally, the multi-channel feature extractor also includes a text feature extractor, and the feature extraction unit 402 is also used to: input the video to be processed into the text feature extractor for text feature extraction to obtain a text feature sequence, and upload it to the cloud side device; when the multimodal processing unit 404 generates the initial input feature according to the second visual feature sequence, it is specifically used to: generate the initial input feature according to the second visual feature sequence and the text feature sequence.

[0129] Further optionally, when the multimodal processing unit 404 generates the initial input feature based on the second visual feature sequence and the text feature sequence, it is specifically used to: encode the second visual feature sequence using the second image encoder on the cloud-side device to obtain a second visual feature embedding vector; encode the text feature sequence using the text encoder on the cloud-side device to obtain a text embedding vector; and splice the second visual feature embedding vector and the text embedding vector to obtain the initial input feature.

[0130] Further optionally, when the video acquisition unit 401 acquires the video to be processed generated by the target end-side device, it is specifically used to: acquire the live video generated by the live broadcast application on the target end-side device during the live broadcast process, as the video to be processed, the live video includes the target product, and the first visual feature sequence and the second visual feature sequence are the visual features of the target product; accordingly, the unimodal processing unit 403 inputs the first visual feature sequence into the unimodal model, and when performing target processing on the video to be processed, it is specifically used to: input the first visual feature sequence into the unimodal model, and perform target detection or content understanding on the live video according to the first visual feature sequence and the third visual feature sequence of the candidate product.

[0131] Further optionally, the feature extraction unit 402 is also used to: pre-extract visual features of the image of the candidate product using the first visual feature extractor and the second visual feature extractor respectively to obtain a third visual feature sequence and a fourth visual feature sequence of the candidate product; establish a correspondence between the third visual feature sequence and the product description information, and provide the correspondence to the unimodal model.

[0132] Further optionally, the feature extraction unit 402 is also used to: upload the third visual feature sequence and the fourth visual feature sequence of the candidate product to the cloud-side device; accordingly, when the multimodal processing unit 404 uses the multimodal model on the cloud-side device to perform target processing on the video to be processed, it is specifically used to: input the first visual feature sequence and the second visual feature sequence into the multimodal model to generate multimodal data corresponding to the target product; generate multimodal data of the candidate product based on the third visual feature sequence and the fourth visual feature sequence of the candidate product and the product description information of the candidate product stored on the cloud-side device; perform target detection or content understanding on the target product based on the multimodal data of the target product and the multimodal data of the candidate product.

[0133] In this embodiment, the first visual feature extractor on the end side can perform personalized feature extraction on the video to obtain a first visual feature sequence, and the second visual feature extractor can extract a second visual feature sequence on the video. The unimodal model on the end side can process the video according to the first visual feature sequence, and the multimodal model on the cloud side can process the video at least according to the first visual feature sequence and the second visual feature sequence when the unimodal model cannot obtain a processing result that meets the requirements. In this way, the first visual feature extractor and the unimodal model on the end side cooperate with each other to complete end-side reasoning and reduce the burden of cloud-side reasoning; the first visual feature extractor has the ability to perform personalized feature extraction on the end-side device, thereby improving the accuracy of end-side reasoning; when end-side reasoning fails, multimodal reasoning can also be performed through the cloud to improve the accuracy of the cloud-side reasoning results. The detailed implementation methods and beneficial effects of each step in the device shown in Figure 4 provided in the embodiment of the present application have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0134] FIG5 is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. As shown in FIG5 , the device includes: a memory 501 and a processor 502 .

[0135] The memory 501 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method used to operate on the electronic device.

[0136] In some embodiments, the processor 502 is coupled to the memory 501 and is used to execute a computer program in the memory 501 to: obtain a video to be processed generated by a target end-side device, where a single-modal model and a multi-channel feature extractor are deployed, and the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor; input the video to be processed into the first visual feature extractor and the second visual feature extractor for visual feature extraction to obtain a first visual feature sequence and a second visual feature sequence; input the first visual feature sequence into the single-modal model to perform target processing on the video to be processed; when the single-modal model cannot obtain a processing result that meets the requirements, upload the first visual feature sequence and the second visual feature sequence to a cloud-side device, and use the multi-modal model on the cloud-side device to perform target processing on the video to be processed.

[0137] In an optional embodiment, the processor 502 is further used to: pre-train the first visual feature extractor using personalized sample data generated by the target end-side device; and pre-train the second visual feature extractor using unified sample data generated by multiple end-side devices.

[0138] In an optional embodiment, when the processor 502 uses the multimodal model on the cloud-side device to perform target processing on the video to be processed, it is specifically used to: input at least the first visual feature sequence into the prompt word generator on the cloud-side device to generate soft prompt words, and provide the generated soft prompt words to the multimodal model; generate initial input features according to the second visual feature sequence, and input the initial input features into the multimodal model; in the multimodal model, embed the soft prompt words into the initial input features to obtain target input features, and perform target processing on the video to be processed according to the target input features.

[0139] In an optional embodiment, when the processor 502 inputs at least the first visual feature sequence into the prompt word generator on the cloud-side device to generate soft prompt words, it is specifically used to: input the first visual feature sequence and the second visual feature sequence into the prompt word generator to generate soft prompt words.

[0140] In an optional embodiment, when the processor 502 inputs the first visual feature sequence and the second visual feature sequence into the prompt word generator to generate soft prompt words, it is specifically used to: encode the first visual feature sequence using a first image encoder to obtain a first visual feature embedding vector; encode the second visual feature sequence using a second image encoder to obtain a second visual feature embedding vector; splice the first visual feature embedding vector and the second visual feature embedding vector, and insert at least one soft prompt marker at the starting position, and insert a segmentation marker at the splicing position to obtain a third feature embedding vector; input the third feature embedding vector into the prompt word generator to generate soft prompt words to obtain at least one soft prompt word corresponding to the at least one soft prompt marker.

[0141] In an optional embodiment, the multi-channel feature extractor also includes a text feature extractor, and the processor 502 is also used to: input the video to be processed into the text feature extractor for text feature extraction to obtain a text feature sequence, and upload it to the cloud side device; when the processor 502 generates the initial input feature according to the second visual feature sequence, it is specifically used to: generate the initial input feature according to the second visual feature sequence and the text feature sequence.

[0142] In an optional embodiment, when the processor 502 generates the initial input feature based on the second visual feature sequence and the text feature sequence, it is specifically used to: encode the second visual feature sequence using the second image encoder on the cloud-side device to obtain a second visual feature embedding vector; encode the text feature sequence using the text encoder on the cloud-side device to obtain a text embedding vector; and concatenate the second visual feature embedding vector and the text embedding vector to obtain the initial input feature.

[0143] In an optional embodiment, when the processor 502 obtains the video to be processed generated by the target end-side device, it is specifically used to: obtain the live video generated by the live broadcast application on the target end-side device during the live broadcast process, as the video to be processed, the live video includes the target product, and the first visual feature sequence and the second visual feature sequence are the visual features of the target product; accordingly, the processor 502 inputs the first visual feature sequence into the single-modal model, and when performing target processing on the video to be processed, it is specifically used to: input the first visual feature sequence into the single-modal model, and perform target detection or content understanding on the live video according to the first visual feature sequence and the third visual feature sequence of the candidate product.

[0144] In an optional embodiment, the processor 502 is further used to: pre-extract visual features from the image of the candidate product using the first visual feature extractor and the second visual feature extractor, respectively, to obtain a third visual feature sequence and a fourth visual feature sequence of the candidate product; establish a correspondence between the third visual feature sequence and the product description information, and provide the correspondence to the unimodal model.

[0145] In an optional embodiment, the processor 502 is also used to: upload the third visual feature sequence and the fourth visual feature sequence of the candidate product to the cloud-side device; accordingly, when the processor 502 uses the multimodal model on the cloud-side device to perform target processing on the video to be processed, it is specifically used to: input the first visual feature sequence and the second visual feature sequence into the multimodal model to generate multimodal data corresponding to the target product; generate multimodal data of the candidate product based on the third visual feature sequence and the fourth visual feature sequence of the candidate product and the product description information of the candidate product stored on the cloud-side device; perform target detection or content understanding on the target product based on the multimodal data of the target product and the multimodal data of the candidate product.

[0146] In other embodiments, the processor 502 is coupled to the memory 501 and is used to execute the computer program in the memory 501 for: obtaining the live video generated by the live broadcast application during the live broadcast, the live broadcast application running on the target end-side device, and the target end-side device is deployed with a single-modal model and a multi-way feature extractor, the multi-way feature extractor including a first visual feature extractor and a second visual feature extractor; inputting the live video into the first visual feature extractor and the second visual feature for visual feature extraction to obtain a first visual feature sequence and a second visual feature sequence; inputting the first visual feature sequence into the single-modal model to perform target detection or content understanding on the target product in the live video; when the single-modal model cannot obtain a processing result that meets the requirements, uploading the first visual feature sequence and the second visual feature sequence to the cloud-side device, and using the multi-modal model on the cloud-side device to perform target detection or content understanding on the target product in the live video. The detailed implementation methods and beneficial effects of each step in the device shown in Figure 5 provided in the embodiment of the present application have been described in detail in the aforementioned embodiment and will not be elaborated here.

[0147] Furthermore, as shown in Figure 5, the electronic device also includes: a communication component 503, a display 504, a power supply component 505, an audio component 506 and other components. Figure 5 only schematically shows some components, which does not mean that the electronic device only includes the components shown in Figure 5. In addition, the components in the dotted box in Figure 5 are optional components, not mandatory components, and the specific components may depend on the product form of the electronic device. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a tablet computer, a smart phone or an IOT device, or it can be a server-side device such as a conventional server, a cloud server or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., it may include the components in the dotted box in Figure 5; if the electronic device of this embodiment is implemented as a server-side device such as a conventional server, a cloud server or a server array, it may not include the components in the dotted box in Figure 5.

[0148] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be executed by the electronic device in the method embodiment shown in FIG. 2 above.

[0149] In this embodiment, the first visual feature extractor on the end side can perform personalized feature extraction on the video to obtain a first visual feature sequence, and the second visual feature extractor can extract a second visual feature sequence of the video. The unimodal model on the end side can process the video based on the first visual feature sequence, and the multimodal model on the cloud side can process the video based on at least the first visual feature sequence and the second visual feature sequence when the unimodal model cannot obtain a processing result that meets the requirements. In this way, the first visual feature extractor and the unimodal model on the end side cooperate with each other to complete end-side reasoning and reduce the burden of cloud-side reasoning; the first visual feature extractor has the ability to perform personalized feature extraction on the end-side device, thereby improving the accuracy of end-side reasoning; when end-side reasoning fails, multimodal reasoning can also be performed through the cloud to improve the accuracy of cloud-side reasoning results.

[0150] The above-mentioned memory can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0151] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0152] The above-mentioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundary of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0153] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.

[0154] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0155] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) that contain computer-usable program code.

[0156] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0157] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0159] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.

[0160] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0161] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0162] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0163] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A video processing system based on end-cloud collaboration, characterized in that: include: The single-modal model and multi-path feature extractor deployed on the edge device, and the multi-modal model deployed on the cloud device; The multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor, which are respectively used to extract a first visual feature sequence and a second visual feature sequence of the video to be processed, and upload them to the cloud-side device; The first visual feature extractor is further used to provide the first visual feature sequence to the unimodal model; the unimodal model is used to perform target processing on the video to be processed according to the first visual feature sequence; The multimodal model is used to perform target processing on the video to be processed at least according to the first visual feature sequence and the second visual feature sequence when the single-modal model cannot obtain a processing result that meets the requirements.

2. The system according to claim 1, characterized in that The first visual feature extractor is trained based on personalized sample data generated by the end-side device where it is located; the second visual feature extractor is trained based on unified sample data generated by multiple end-side devices.

3. The system according to claim 1, characterized in that Also includes: A prompt word generator deployed on the cloud-side device, used to generate soft prompt words at least according to the first visual feature sequence, and output the soft prompt words to the multimodal model; The multimodal model is specifically used to: when the single-modal model cannot obtain a processing result that meets the requirements, generate an initial input feature according to the second visual feature sequence, and embed the soft prompt word into the initial input feature to obtain a target input feature; The video to be processed is subjected to target processing according to the target input feature.

4. The system according to claim 3, characterized in that The cloud-side device is also deployed with: A first image encoder, used for encoding the first visual feature sequence to obtain a first visual feature embedding vector; a second image encoder, configured to encode the second visual feature sequence to obtain a second visual feature embedding vector; The first splicing module is used to splice the first visual feature embedding vector and the second visual feature embedding vector, insert at least one soft prompt mark at the starting position, and insert a segmentation mark at the splicing position to obtain a third feature embedding vector, and input the third feature embedding vector into the prompt word generator to generate a soft prompt word.

5. The system according to claim 3, characterized in that The multi-channel feature extractor also includes: a text feature extractor, which is used to extract a text feature sequence of the video to be processed and upload it to the cloud-side device; When generating the initial input features according to the second visual feature sequence, the multimodal model is specifically used to: generate the initial input features according to the second visual feature sequence and a text feature sequence.

6. The system according to claim 5, characterized in that The cloud-side device is also deployed with: a second image encoder, configured to encode the second visual feature sequence to obtain a second visual feature embedding vector; A text encoder, used for encoding the text feature sequence to obtain a text embedding vector; The second concatenation module is used to concatenate the second visual feature embedding vector and the text embedding vector to obtain the initial input feature.

7. A video processing method based on end-cloud collaboration, characterized in that: include: Acquire a to-be-processed video generated by a target end-side device, wherein a single-modal model and a multi-channel feature extractor are deployed on the target end-side device, and the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor; Inputting the video to be processed into the first visual feature extractor and the second visual feature extractor to extract visual features, so as to obtain a first visual feature sequence and a second visual feature sequence; Inputting the first visual feature sequence into the unimodal model, and performing target processing on the video to be processed; When the unimodal model cannot obtain a processing result that meets the requirements, the first visual feature sequence and the second visual feature sequence are uploaded to a cloud-side device, and the multimodal model on the cloud-side device is used to perform target processing on the video to be processed.

8. The method according to claim 7, characterized in that Utilizing the multimodal model on the cloud-side device to perform target processing on the video to be processed includes: at least inputting the first visual feature sequence into a prompt word generator on the cloud-side device to generate soft prompt words, and providing the generated soft prompt words to the multimodal model; generating initial input features according to the second visual feature sequence, and inputting the initial input features into the multimodal model; In the multimodal model, the soft prompt word is embedded in the initial input feature to obtain a target input feature, and the to-be-processed video is subjected to target processing according to the target input feature.

9. The method according to claim 8, characterized in that Inputting at least the first visual feature sequence into a prompt word generator on the cloud-side device to generate a soft prompt word includes: Encoding the first visual feature sequence using a first image encoder to obtain a first visual feature embedding vector; Encoding the second visual feature sequence using a second image encoder to obtain a second visual feature embedding vector; splicing the first visual feature embedding vector and the second visual feature embedding vector, inserting at least one soft prompt marker at a starting position, and inserting a segmentation marker at a splicing position, so as to obtain a third feature embedding vector; The third feature embedding vector is input into a prompt word generator to generate a soft prompt word, so as to obtain at least one soft prompt word corresponding to the at least one soft prompt mark.

10. The method according to claim 8, characterized in that The multi-path feature extractor also includes a text feature extractor, and the method further includes: Inputting the video to be processed into the text feature extractor to extract text features to obtain a text feature sequence, and uploading it to the cloud-side device; Generating initial input features according to the second visual feature sequence includes: generating the initial input features according to the second visual feature sequence and a text feature sequence.

11. The method according to claim 10, characterized in that Generating the initial input feature according to the second visual feature sequence and the text feature sequence includes: Encoding the second visual feature sequence using a second image encoder on the cloud-side device to obtain a second visual feature embedding vector; Encoding the text feature sequence using a text encoder on the cloud-side device to obtain a text embedding vector; The second visual feature embedding vector and the text embedding vector are concatenated to obtain the initial input feature.

12. The method according to any one of claims 7 to 11, characterized in that: Obtain the video to be processed generated by the target end device, including: Acquire a live video generated by a live broadcast application on the target end-side device during a live broadcast process as a video to be processed, wherein the live video includes a target product, and the first visual feature sequence and the second visual feature sequence are visual features of the target product; Accordingly, the first visual feature sequence is input into the unimodal model, and target processing is performed on the video to be processed, including: The first visual feature sequence is input into the unimodal model, and target detection or content understanding is performed on the live video according to the first visual feature sequence and the third visual feature sequence of the candidate product.

13. A video processing method based on end-cloud collaboration, characterized in that: include: Acquire a live video generated by a live broadcast application in a live broadcast, wherein the live broadcast application runs on a target end-side device, and a single modal model and a multi-channel feature extractor are deployed on the target end-side device, wherein the multi-channel feature extractor includes a first visual feature extractor and a second visual feature extractor; Inputting the live video into the first visual feature extractor and the second visual feature extractor to extract visual features, so as to obtain a first visual feature sequence and a second visual feature sequence; Inputting the first visual feature sequence into the unimodal model to perform target detection or content understanding on the target product in the live video; When the unimodal model cannot obtain processing results that meet the requirements, the first visual feature sequence and the second visual feature sequence are uploaded to the cloud-side device, and the multimodal model on the cloud-side device is used to perform target detection or content understanding on the target product in the live video.

14. An electronic device, characterized in that: include: A memory and a processor; wherein the memory is used to: store one or more computer instructions; and the processor is used to execute the one or more computer instructions to: execute the steps in the method described in any one of claims 7-13.

15. A computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 7 to 13.

Citation Information

Patent Citations

  • All-media news intelligent cataloguing method based on multi-modal information fusion understanding

    CN112818906A

  • Video frame processing method and device, electronic equipment, storage medium and program product

    CN115147754A

  • Multi-modal knowledge question-answering method and system for 5G message

    CN116932731A

  • Video understanding method and device

    CN116994171A

  • Video classification method, device and equipment and computer readable storage medium

    CN117036833A

Cited By

  • Multi-modal interactive video analysis method and device, computer equipment and storage medium

    CN122135178A