Video processing method, electronic device, and computer storage medium

By directly acquiring and processing video streams within the application, the time-consuming detection issue caused by manually triggering video capture functions is resolved, enabling rapid target identification and data retrieval, and improving the user experience.

CN116385916BActive Publication Date: 2026-04-10ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the video recording function needs to be manually triggered to enter the video recording state, which results in a long detection time and affects the user experience.

Method used

By launching a video stream-based retrieval task within the application, the video stream containing the target subject can be directly obtained for subject identification and data retrieval, avoiding the need to manually trigger the video capture function.

Benefits of technology

It enables rapid identification of target entities, reduces detection time, and improves interaction fluency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385916B_ABST
    Figure CN116385916B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video processing method, an electronic device and a computer storage medium, wherein the method comprises: determining a video stream-based retrieval task started in an application program, that is, obtaining a generated video stream containing a target subject when the application program is running, and then directly performing subject recognition based on the generated video stream, without manually triggering a video shooting function to enter a video shooting state to generate a video stream, so that the target subject can be quickly identified, and further data associated with the target subject is retrieved based on the target subject, thereby shortening the time consumption of the detection scheme, improving the smoothness of interaction, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computer technology, and in particular, to a video processing method, an electronic device and a computer storage medium. BACKGROUND

[0002] In a target subject detection scheme, after starting a camera device, a video shooting function needs to be manually triggered to enter a video shooting state, and then a video stream is obtained, and image processing is performed on the video stream to detect a target subject in the video stream. Therefore, when applied in some scenes, since the video shooting function needs to be manually triggered to enter the video shooting state, the time consumption of the entire detection scheme is relatively long, the smoothness of interaction is affected, and thus the user experience is reduced. SUMMARY

[0003] Therefore, embodiments of the present application provide a video processing scheme to at least partially solve the above problems.

[0004] According to a first aspect of embodiments of the present application, a video processing method is provided, comprising:

[0005] determining a video stream-based retrieval task started in an application program, and obtaining a generated video stream containing the target subject when the application program is running;

[0006] performing subject identification processing on the video stream to identify a target subject in the video stream;

[0007] based on the target subject identified from the video stream, performing the retrieval task to retrieve data associated with the target subject, the data including a recommended object similar to the target subject and / or attribute description data of the recommended object.

[0008] According to a second aspect of embodiments of the present application, an electronic device is provided, comprising a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method of the first aspect.

[0009] According to a third aspect of embodiments of the present application, a computer storage medium is provided, and the computer storage medium stores a computer program, and the program is executed by a processor to implement the method of the first aspect.

[0010] According to a fourth aspect of embodiments of the present application, a computer program product is provided, comprising computer instructions, and the computer instructions instruct a computing device to perform operations corresponding to the method of the first aspect.

[0011] According to the video processing scheme provided in the embodiments of the present application, the video stream-based retrieval task started in the application program is determined, that is, the video stream generated based on the target subject can be obtained when the application program is run, and then the subject recognition is directly performed based on the generated video stream, without manually triggering the video shooting function to enter the video shooting state to generate the video stream, so that the target subject can be quickly recognized, and further, the data associated with the target subject is retrieved based on the target subject, thereby shortening the time consumption of the detection scheme, improving the smoothness of the interaction, and improving the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0013] Figure 1 An exemplary system suitable for the video processing method according to the embodiments of the present application is shown.

[0014] Figure 2A A flowchart of a video processing method is shown.

[0015] Figure 2B A flowchart of a subject recognition process is shown.

[0016] Figure 3A An application scenario of a video processing method according to the embodiments of the present application is shown.

[0017] Figure 3B An application scenario of a multi-thread control video processing method according to the embodiments of the present application is shown.

[0018] Figure 4 A structural diagram of an electronic device according to the embodiments of the present application is shown. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those skilled in the art should belong to the scope of protection of the embodiments of the present application.

[0020] The specific implementation of the embodiments of the present application will be further described below with reference to the drawings of the embodiments of the present application.

[0021] In the technical solutions provided by the embodiments of the present application, the target subject is directly identified based on the generated video stream containing the target subject when the application program is running, without manually triggering the video shooting function to enter the video shooting state to generate the video stream. Meanwhile, the operation of identifying the target subject can be performed as soon as the generated video stream containing the target subject is obtained, thereby simplifying the process of the detection scheme, shortening the time consumption of the detection scheme, improving the smoothness of the interaction, improving the user experience, and achieving the effect of seeing and searching.

[0022] Figure 1 An exemplary system suitable for the video processing method according to the embodiments of the present application is shown. As shown in the figure, Figure 1 The system 100 can include a cloud server 102, a communication network 104 and / or one or more user devices 106, Figure 1 The user devices are exemplified as multiple user devices.

[0023] The cloud server 102 can be any appropriate device for storing information, data, application programs and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the video processing method provided by the embodiments of the present application can be integrated in an application program and stored on the cloud server 102.

[0024] In some embodiments, the communication network 104 can be any appropriate combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the Internet, an intranet, a wide-area network (WAN), a local-area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user devices 106 can connect to the communication network 104 through one or more communication links, such as communication links 112, which can link to the cloud server 102 via one or more communication links, such as communication links 114. The communication links can be any communication links suitable for communicating data among the user devices 106 and the cloud server 102, such as network links, dial-up links, wireless links, hard-wired links, any other suitable communication links, or any suitable combination of such links.

[0025] The user devices 106 download the application program from the communication network to the local to execute the video processing method provided by the embodiments of the present application locally on the user devices.

[0026] In some embodiments, user equipment 106 may include any suitable type of device. For example, in some embodiments, user equipment 106 may include mobile devices, tablet computers, laptop computers, desktop computers, wearable computers, game consoles, media players, vehicle entertainment systems, and / or any other suitable type of user equipment.

[0027] Referring to the above system, this application provides a video processing method, which will be described below through several embodiments.

[0028] Of course, it should be noted here that... Figure 1 In this embodiment, the video processing method is illustrated by executing it locally on the user device. However, this does not exclusively limit the execution to the user device. In fact, in some application scenarios, it can also be executed on a cloud server, and the results can then be pushed to the user device.

[0029] Based on the above system, this application provides a video processing method, which will be described below through several embodiments.

[0030] Figure 2A A flowchart illustrating a video processing method is shown. For example... Figure 2A As shown, in this embodiment, the video processing method includes the following steps S201-S203:

[0031] S201. Determine the video stream-based retrieval task initiated in the application, and obtain the video stream containing the target subject generated when the application is running.

[0032] In some examples, the application is not particularly limited and may include any application that integrates video stream-based retrieval, such as applications that provide consumer services.

[0033] In some examples, as previously mentioned, the application runs on an electronic device that may be equipped with a camera, and the video stream may specifically be captured by that camera.

[0034] In some examples, to capture the video stream, the main thread of the electronic device can be started to initialize the camera, and a camera configuration thread can be started to activate the camera. Then, a camera callback thread can be started to control the camera to point at the target subject to generate the video stream. It should be noted that by pointing the camera at the target subject, the corresponding video stream is generated in real time, without the need for manual camera operation, such as clicking a video capture button.

[0035] In some examples, the condition for starting the video stream-based retrieval task in the application can be set according to the requirements of the application scenario, and is not uniquely limited as long as the condition for running the video processing method of the embodiment of the present application is met.

[0036] S202, subject recognition processing is performed on the video stream to identify a target subject in the video stream.

[0037] In some examples, step S202 can be started by a configured subject recognition processing thread, which can be executed after the camera callback thread successfully completes the generation of the video stream.

[0038] In some examples, the subject recognition processing on the video stream to identify a target subject in the video stream in step S202 can include subject recognition processing on the video stream in units of video frames to identify a target subject in the video stream. Subject recognition processing in units of video frames is equivalent to dividing the task of subject recognition processing, so that real-time identification processing of the target subject that may appear in the video stream can be performed. That is, the video stream is generated in real time, and the subject recognition processing is also performed in real time along with the real-time generation of the video stream, or it is also called online identification processing.

[0039] In some examples, as shown in Figure 2B A flowchart of subject recognition processing is shown. Wherein, the subject recognition processing on the video stream to identify a target subject in the video stream includes:

[0040] S212, in units of video frames, each video frame is converted to generate a corresponding video frame image.

[0041] S222, subject recognition processing is performed on the video stream to identify the target subject therein.

[0042] The conversion processing includes scaling and / or flipping of the picture. By scaling the picture, the size of the video frame image can be reduced, the amount of data participating in subsequent subject recognition processing can be reduced, and the efficiency of subject recognition processing can be improved. For picture rotation, for example, when the camera is in a vertical screen shooting state and the video frame image is displayed in a vertical screen direction, the video frame image is converted from a vertical screen to a horizontal screen for subject recognition processing, which requires the video frame image to be displayed in a horizontal screen direction.

[0043] Of course, in some examples, if the video frame image just meets the requirements of the subject recognition processing for the video frame image, such as size, display direction, etc., the conversion processing process can be omitted.

[0044] It should be noted that the conversion between the portrait screen and the landscape screen is merely an example and is not the only one.

[0045] Of course, in some other examples, the subject recognition processing can be performed on the entire video stream after generating the video stream for a period of time after aligning with the target subject. Unlike the real-time processing described above, the subject recognition processing on the video stream is offline recognition processing.

[0046] To this end, the entire video stream can be directly converted into a series of video frame images by referring to the processing described above, and the subject recognition processing can be performed on the series of video frame images to identify the target subject in the video stream.

[0047] In some examples, the subject recognition processing on the video stream to identify the target subject therein includes subject recognition processing on the video stream based on a first subject recognition model that has completed initialization, and subject recognition processing on the video stream based on a second subject recognition model as a backup if the first subject recognition model has not completed initialization.

[0048] In some examples, the first subject recognition model can be a neural network model that has completed training. The first subject recognition model can be trained by using sample pictures collected in advance to meet the requirements of the embodiments of the present application.

[0049] In some examples, the second subject recognition model can be a lightweight picture recognition model.

[0050] In some examples, whether the first subject recognition model has completed initialization can be determined before the subject recognition processing on the video stream to identify the target subject therein is performed in units of video frames. If the first subject recognition model has completed initialization, each frame of video can be converted into a corresponding video frame image in units of video frames, and the subject recognition processing on the video stream can be performed based on the first subject recognition model that has completed initialization to identify the target subject therein. Otherwise, the second subject recognition model as a backup can be switched to, each frame of video can be converted into a corresponding video frame image in units of video frames, and the subject recognition processing on the video stream can be performed based on the second subject recognition model to identify the target subject therein.

[0051] In some examples, the first subject identification model can be switched to the second subject identification model for subject identification processing when the initialization of the first subject identification model is not completed, so as to ensure that the target subject can still be identified. For example, when the initialization of the first subject identification model is not completed, the second subject identification model can be switched to the second subject identification model for subject identification processing through a bottom link.

[0052] In some examples, the interval for conversion processing of adjacent two video frames is a first interframe interval. This step can be performed between step S201 and step S202.

[0053] The size of the first interframe interval can be determined according to the requirements of the conversion processing, such as 0.05s in some embodiments, as long as it can meet the conversion processing of the video stream frame by frame, so that the conversion processing is completed for a frame of video, and then the conversion processing is performed for the next frame of video, avoiding the conflict of conversion processing when the conversion processing is performed on two continuous video frames.

[0054] In some examples, the method further includes: the interval for subject identification processing of video frame images corresponding to adjacent two video frames is a second interframe interval.

[0055] The size of the second interframe interval can be determined according to the requirements of the subject identification processing, such as 2s in some examples, as long as it can meet the subject identification processing of the video frame images corresponding to the video stream frame by frame, so that the subject identification processing is completed for a video frame image of a frame of video, and then the subject identification processing is performed for the video frame image corresponding to the next frame of video, avoiding the conflict of subject identification processing when the subject identification processing is performed on the video frame images of two continuous video frames.

[0056] S203, based on the target subject identified from the video stream, performing the retrieval task to retrieve data associated with the target subject.

[0057] In some examples, a retrieval database can be established, and the database includes a plurality of data, such as pictures of recommended objects having the same or attributes as the target subject, or attribute description data of recommended objects.

[0058] In some examples, the association degree between the target subject and the data in the database can be calculated by an association degree calculation method, such as similarity, and the data with the largest association degree is taken as the data associated with the target subject. The association can be defined according to the application scenario.

[0059] In some examples, the method further includes: generating a search result status code according to the search result; and generating a search result operation code according to the search result status code, the search result status code being used for utility classification of the search result, and the search result operation code being used for operation classification of the search result.

[0060] The utility classification of the search result may include, for example, that the video stream meets the requirement in terms of resolution but no valid data is searched, that the video stream does not meet the requirement in terms of resolution but valid data is searched, and the like. Different search result status codes may be represented by different digital codes.

[0061] In some examples, the search result operation code may be generated by performing statistical analysis on the search result represented by the search result status code and assigning a value to the search result operation code by digital coding. For example, the search result operation code may be classified into three categories: the video stream meets the requirement in terms of resolution and valid data is searched; the video stream does not meet the requirement in terms of resolution and / or valid data is not searched; and the video stream meets the requirement in terms of resolution but it is uncertain whether valid data is searched.

[0062] Further, the next operation may be performed according to the search result operation code, such as re-performing subject identification processing, or ignoring the search result, or displaying the search result.

[0063] To this end, the method may further include: generating a search result display card according to the data associated with the target subject in response to the search result operation code.

[0064] If the target subject is identified based on the video frame, when the video frame meets the requirement in terms of resolution and valid data is searched, the attribute description data of the recommended object associated with the target subject is loaded on the search result display card for display. For example, the price and origin of the target subject.

[0065] When there are multiple generated search result display cards, different search result display cards may be displayed at different positions on the content interface of the video stream, and the confidence of the searched data may be displayed at the same time.

[0066] In some examples, the method further includes: in response to the search result operation code, obtaining and displaying the search result generated according to the data associated with the target subject according to the pulled search result page task.

[0067] Exemplarily, the cloud server performs statistical analysis on the data associated with the target subject to generate a search result, and based on the search result, the electronic device displays the search result.

[0068] In some examples, on the basis of the above embodiment, a predetermined pre-detection can be further included to determine whether the condition for starting the video stream-based search task in the application is met, and the pre-detection includes a plurality of detection items, and the detection items include at least one of scene type detection, state type detection, video stream stability detection, and resource occupation type detection. This step is performed in step S201.

[0069] For example, the pre-detection is implemented by monitoring the camera callback thread. When the detection result meets the requirement, the subject recognition processing thread is jumped to.

[0070] By performing the predetermined pre-detection, it is ensured that the video processing method of the present application is started only when the condition for starting the video stream-based search task is met, or in other words, the dirty data is effectively filtered out, and the smooth execution of the subsequent video processing scheme is effectively ensured, and the search efficiency is improved.

[0071] In some examples, the scene type detection includes detecting whether the pop-up window business of the application is completed, and whether the target subject is aligned when the application is running. If the detection results of both are yes, it is determined that the current scene meets the condition for starting the video stream-based search task in the application. The pop-up window business is, for example, the use instruction of the application. If the camera shooting function is triggered when the application is running, or the picture is selected from the local of the electronic device, the normal picture-based target subject search scheme is directly executed. When the target subject is aligned, the camera can generate the video stream.

[0072] In some examples, the state type detection includes detecting whether the plurality of detection items are executed according to the predetermined detection order. If yes, it is determined that the current state meets the condition for starting the video stream-based search task in the application.

[0073] Exemplarily, the detection order of the detection items can be determined according to the application scenario. By formulating the detection order, only when the camera callback thread is monitored and it is determined that the detection items all meet the condition for starting the video stream-based search task in the application, the subject recognition processing thread is jumped to perform the subject recognition processing on the video stream to identify the target subject in the video stream.

[0074] Exemplarily, the detection of the detection items can be implemented by a state machine.

[0075] In some examples, the video stream stability detection comprises: detecting whether the definition of the video stream meets a definition requirement, whether the illumination of the video stream meets an illumination intensity requirement, and whether the electronic device running the application program meets a hardware requirement; if the results of the three determinations are all yes, it is determined that the stability of the video stream meets the requirement of starting the video stream-based retrieval task in the application program.

[0076] For example, the definition of the video stream can be detected by frame-by-frame detection, for example, by comparing the definition of the video stream with a set definition threshold, and if the definition is greater than the definition threshold, it is considered that the definition requirement is met, otherwise, it is considered that the definition requirement is not met. Alternatively, in some other examples, the difference between the hash values of the motion data of two adjacent video frames is calculated, and if the difference is less than a set hash value threshold, it is considered that the definition requirement is met, otherwise, it is considered that the definition requirement is not met.

[0077] In some examples, the resource occupation detection comprises: detecting whether the video stream is not involved in other non-body recognition processing, and whether the remaining memory of the electronic device running the application program meets the requirement of performing the body recognition processing; if the results of the two determinations are both yes, it is determined that the resource occupation meets the requirement of starting the video stream-based retrieval task in the application program.

[0078] By detecting whether the video stream is not involved in other non-body recognition processing, the priority of video stream consumption is set. When the video stream is not involved in other non-body recognition processing, the video stream can be involved in the body recognition processing of the embodiment of the application. In addition, by detecting the remaining memory, the body recognition processing is performed when the remaining memory supports the body recognition processing, thereby avoiding memory overflow.

[0079] Here, it should be noted that the video frame of the above embodiment of the application can be based on YUV color space and RGB color space, and therefore the specific format thereof can be determined according to the requirements of the application scenario.

[0080] Figure 3A An application scenario of a video processing method according to an embodiment of the application is shown. As shown in FIG. 1, the application scenario comprises an electronic device 100 and a video stream 200. Figure 3AAs shown, in this embodiment, the target subject is a schoolbag, and a video stream can be generated by aligning the target object through the camera on the electronic device. The user is reminded on the electronic device to keep the mobile phone stable and align the product. In units of video frames, subject recognition processing is performed on the video stream to identify the schoolbag in the video stream, and based on the schoolbag being identified, a database is searched to obtain attribute description data associated with the schoolbag, such as name, evaluation, price, etc., which is displayed in the form of a card in the video frame. At the same time, a number of schoolbags with a certain similarity to the schoolbag as the target subject are searched.

[0081] Figure 3B An application scenario of the video processing method based on multi-thread control according to an embodiment of the present application is shown. Corresponding to the above Figure 3A , in order to implement Figure 3A the video processing method, multi-thread control is configured. Specifically, the main thread of the electronic device initializes the camera, and then calls the camera configuration thread to start the camera. The pre-detection recorded in the above embodiment is performed in the camera callback thread, and after the pre-detection passes, the video stream generated by the camera aligning the schoolbag is frame-callbacked. At this time, it is judged whether the camera photographing is triggered or not. If yes, it jumps to the picture processing to directly perform the subject recognition processing on the photographed picture. If no, the video stream is detected in the video frame, the corresponding video frame image is further generated, and the flow of calling the subject recognition processing is started. The subject recognition thread callback is entered, the subject recognition processing is called back, that is, the subject recognition processing is performed on each frame of video corresponding to the video frame image, the data associated with the target subject is searched, such as the schoolbag attribute description data and the similar schoolbags with a certain similarity to the schoolbag, and it is further judged whether the debugging mode is started or not. If yes, the debugging processing is performed, such as saving the video frame image corresponding to the current video frame, the context information of the subject recognition processing, etc., so as to debug the subject recognition processing. If no, it jumps to the random sub-thread to obtain the result of the subject recognition processing, and the retrieval result page is pulled up by the main thread to display and show the retrieval result generated according to the data associated with the target subject.

[0082] An embodiment of the present application further provides a video processing device, comprising:

[0083] A data acquisition unit is configured to determine a video stream-based retrieval task started in an application program, and acquire a generated video stream containing a target subject when the application program is running.

[0084] An identification unit is configured to perform subject recognition processing on the video stream to identify the target subject in the video stream.

[0085] The task execution unit is configured to execute the search task based on the target subject identified from the video stream to search for data associated with the target subject, the data including recommended objects similar to the target subject and / or attribute description data of the recommended objects.

[0086] The identification unit is configured to perform subject identification processing on the video stream based on the first subject identification model that has completed initialization to identify the target subject in the video stream.

[0087] If the initialization of the first subject identification model is not completed, the identification unit is configured to perform subject identification processing on the video stream based on a second subject identification model that is ready for use to identify the target subject in the video stream.

[0088] The identification unit is configured to perform conversion processing on each frame of video in units of video frames to generate corresponding video frame images.

[0089] The identification unit is configured to perform subject identification processing on the video frame images to identify the target subject in the corresponding video frames.

[0090] The identification unit is configured to perform subject identification processing on the video frames in the video stream to identify the target subject in the video frames, wherein the time interval for performing the subject identification processing on adjacent two video frames is a second interframe interval.

[0091] The apparatus further includes a search result processing unit configured to:

[0092] generate a search result status code based on the search result, the search result status code being used for utility classification of the search result; and

[0093] generate a search result operation code based on the search result status code, the search result operation code being used for operation classification of the search result.

[0094] The search result processing unit is configured to generate a search result presentation card based on the data associated with the target subject in response to the search result operation code.

[0095] The search result processing unit is configured to obtain and display the search result generated based on the data associated with the target subject in response to the search result operation code according to a pulled search result page task.

[0096] Exemplarily, the device further comprises a pre-detection unit configured to perform predetermined pre-detection to determine whether a condition for starting the video stream-based search task in the application is met, wherein the pre-detection comprises a plurality of detection items, and the detection items comprise at least one of scene detection, state detection, video stream stability detection, and resource occupation detection.

[0097] Exemplarily, the scene detection comprises: detecting whether a pop-up window service of the application is completed, and whether the target subject is aligned when the application is running; if the detection results of both are yes, it is determined that the current scene meets the condition for starting the video stream-based search task in the application.

[0098] Exemplarily, the state detection comprises: detecting whether the plurality of detection items are completed in a predetermined detection order; if yes, it is determined that the current state meets the condition for starting the video stream-based search task in the application.

[0099] Exemplarily, the video stream stability detection comprises: detecting whether a definition of the video stream meets a definition requirement, whether an illumination of the video stream meets an illumination intensity requirement, and whether an electronic device running the application meets a hardware requirement; if the detection results of all the three are yes, it is determined that the stability of the video stream meets the condition for starting the video stream-based search task in the application.

[0100] Exemplarily, the resource occupation detection comprises: detecting whether the video stream is not involved in other non-subject recognition processing, and whether a remaining memory of the electronic device running the application meets a requirement for the subject recognition processing; if the detection results of both are yes, it is determined that the resource occupation meets the condition for starting the video stream-based search task in the application.

[0101] Referring to Figure 4 , a structural schematic diagram of an electronic device according to an embodiment of the present application is shown, and the embodiment of the present application does not limit the specific implementation of the electronic device.

[0102] As Figure 4 shown, the electronic device can comprise a processor 402, a communications interface 404, a memory 406, and a communications bus 408.

[0103] Among them:

[0104] The processor 402, the communications interface 404, and the memory 406 complete mutual communication through the communications bus 408.

[0105] The communications interface 404 is configured to communicate with other electronic devices or servers.

[0106] The processor 402 is configured to execute the program 410, and particularly configured to execute the steps of the above-mentioned video processing method embodiments.

[0107] Specifically, the program 410 can include program codes including computer operation instructions.

[0108] The processor 402 can be a CPU, or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device can be the same type of processor, such as one or more CPUs; or can be different types of processors, such as one or more CPUs and one or more ASICs.

[0109] The memory 406 is configured to store the program 410. The memory 406 can include a high-speed RAM memory, and can also include a non-volatile memory such as at least one disk memory.

[0110] The program 410 can be specifically configured to cause the processor 402 to perform the operations corresponding to the video processing described in any of the above-mentioned method embodiments.

[0111] The specific implementation of each step in the program 410 can refer to the corresponding description in the above-mentioned method embodiments and has corresponding beneficial effects, which will not be described here. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned device and module can refer to the corresponding process description in the above-mentioned method embodiments, which will not be described here.

[0112] The embodiments of the present application also provide a computer program product, including computer instructions, which instruct a computing device to perform the operations corresponding to any of the above-mentioned video processing methods.

[0113] The embodiments of the present application also provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement any of the above-mentioned video processing methods.

[0114] It should be noted that, according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or part of the operations of the components / steps can be combined into a new component / step, to achieve the purpose of the embodiments of the present application.

[0115] The methods according to the embodiments of the present application described above can be implemented in hardware, firmware, or software, or a combination of them, and can be stored in a recording medium such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk, or be downloaded from a network originally stored in a remote recording medium or a non-transitory machine-readable medium and stored in a local recording medium, so that the methods described herein can be processed by such software on a recording medium using a general-purpose computer, a special-purpose processor, or programmable or special-purpose hardware such as an ASIC or an FPGA. It can be understood that the computer, processor, microprocessor controller, or programmable hardware includes a storage component (for example, RAM, ROM, flash memory, etc.) that can store or receive software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code for implementing the methods shown herein, the execution of the code will convert the general-purpose computer into a special-purpose computer for executing the methods shown herein.

[0116] Those skilled in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.

[0117] The above embodiments are only used to illustrate but not limit the embodiments of the present application, and a person of ordinary skill in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application, therefore all equivalent technical solutions belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application should be defined by the claims.

Claims

1. A method for video processing, comprising: determining a video stream-based retrieval task initiated in an application, and obtaining a generated video stream containing a target subject when the application is running on an electronic device, wherein the video stream is generated by starting a main thread of the electronic device to complete initialization of a camera configured on the electronic device, starting a camera configuration thread to turn on a camera, and starting a camera callback thread to control the camera to aim at a target subject to generate the video stream; performing subject identification processing on the video stream to identify the target subject in the video stream; based on the target subject identified from the video stream, performing the retrieval task to retrieve data associated with the target subject, the data including a recommended object similar to the target subject and / or attribute description data of the recommended object.

2. The method of claim 1, wherein, The subject identification processing on the video frame image to identify the target subject in the corresponding video frame includes performing subject identification processing on the video stream based on a first subject identification model that has completed initialization to identify the target subject in the video stream. If the initialization of the first subject identification model is not completed, performing subject identification processing on the video stream based on a second subject identification model that is a backup to identify the target subject in the video stream.

3. The method of claim 1, wherein, The subject identification processing on the video stream to identify the target subject in the video stream includes: performing conversion processing on each frame of video in units of video frames to generate a corresponding video frame image; performing subject identification processing on the video frame image to identify the target subject in the corresponding video frame, wherein a time interval for performing conversion processing on two adjacent video frames is a first inter-frame interval.

4. The method of claim 1, wherein, The subject identification processing on the video stream to identify the target subject in the video stream includes performing subject identification processing on video frames in the video stream to identify the target subject in the video frames, wherein a time interval for performing the subject identification processing on two adjacent video frames is a second inter-frame interval.

5. The method of claim 1, wherein, The method further comprises: generating a retrieval result status code based on the retrieval result, the retrieval result status code being used for utility classification of the retrieval result; generating a retrieval result operation code based on the retrieval result status code, the retrieval result operation code being used for operation classification of the retrieval result.

6. The method of claim 5, wherein, The method further comprises: in response to the retrieval result operation code, generating a retrieval result presentation card based on the data associated with the target subject.

7. The method of claim 6, wherein, The method further comprises: in response to the retrieval result operation code, obtaining and displaying the retrieval result generated based on the data associated with the target subject according to a pulled retrieval result page task.

8. The method according to any one of claims 1 to 7, wherein, The method further comprises: performing a predetermined pre-detection to determine whether a condition for starting the video stream-based retrieval task in the application is met, the pre-detection comprising a plurality of detection items, the detection items comprising at least one of a scene type detection, a state type detection, a video stream stability detection, and a resource occupation type detection.

9. The method of claim 8, wherein, The scene type detection comprises: detecting whether a pop-up window service of the application is completed, and whether the target subject is in the field of view during running of the application; and if both are yes, determining that a current scene meets the condition for starting the video stream-based retrieval task in the application.

10. The method of claim 8, wherein, The state type detection comprises: detecting whether the plurality of detection items are completed in a predetermined detection order; and if yes, determining that a current state meets the condition for starting the video stream-based retrieval task in the application.

11. The method of claim 8, wherein, The video stream stability detection comprises: detecting whether a definition of the video stream meets a definition requirement, whether an illumination of the video stream meets an illumination intensity requirement, and whether an electronic device running the application meets a hardware requirement; and if all are yes, determining that a stability of the video stream meets the condition for starting the video stream-based retrieval task in the application.

12. The method of claim 8, wherein, The resource occupation type detection comprises: detecting whether the video stream is not involved in other non-subject identification processing, and whether a remaining memory of the electronic device running the application meets a requirement for the subject identification processing; and if both are yes, determining that the resource occupation meets the condition for starting the video stream-based retrieval task in the application.

13. An electronic device comprising: A processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface being in communication with each other through the communication bus; The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method of any one of claims 1-12.

14. A computer storage medium having a computer program stored thereon, the program being executed by a processor to implement the method of any one of claims 1-12.

15. A computer program product comprising computer instructions for instructing a computing device to perform operations corresponding to the method of any one of claims 1-12.

Citation Information

Patent Citations

  • Video-based commodity recommendation method and device

    CN106055710A

  • Image information processing method, system and device and storage medium

    CN114419699A