Commodity pushing method and device and storage medium
By receiving user text information during video playback for item identification and product information retrieval, and directly pushing target products on the video interface, solving the problem of jumping during video viewing affects integrity, and achieving low-cost product information acquisition.
Patent Information
- Application Number
- CN202410060391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
When users have the intention to purchase when watching videos, they need to jump to the shopping platform for product search, which affects the viewing integrity of the video content and increases the cost of obtaining product information.
During the video playback process, receive the target text information entered by the user, identify the item, obtain the target item information, and search in the preset product information library, push the target product information to the video playback interface to avoid jumping to the shopping platform.
Without affecting the integrity of video viewing, the cost of users to obtain product information that meets purchase preferences and needs is reduced, and shopping satisfaction is improved.
Smart Images

Figure CN120338905A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and in particular, to a method and device for pushing commodities and a storage medium. Background Art
[0002] As the audience of videos continues to increase, more and more merchants choose to promote their commodities to users through videos. For example, they can promote their commodities to users by having the main characters in the videos wear their commodities. When a user has the intention to purchase a commodity while watching a video, they can search for and purchase the same commodity through a shopping platform integrated with commodity retrieval and recommendation functions.
[0003] However, when a user has the intention to purchase a commodity while watching a video, they need to jump the current video playback interface to the shopping platform for commodity retrieval in order to find relevant commodity information. Therefore, throughout the process, it will not only affect the integrity of the user's viewing of the video content, but also increase the cost for the user to obtain commodity information. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of this application provide a method and device for pushing commodities and a storage medium, which can reduce the cost for the user to obtain commodity information without affecting the integrity of the user's viewing of the video content.
[0006] On the one hand, embodiments of this application provide a method for pushing commodities, including the following steps:
[0007] During video playback, receive the target text information input by a first object, and perform item recognition on the current video playback content according to the target text information to obtain target item information;
[0008] Retrieve commodity information in a pre-set commodity information library according to the target item information to obtain the commodity information of the target commodity;
[0009] Push the commodity information of the target commodity to the first object on the screen interface of the current video playback.
[0010] On the other hand, embodiments of this application also provide a device for pushing commodities, including:
[0011] An item recognition unit, configured to receive the target text information input by a first object during video playback, and perform item recognition on the current video playback content according to the target text information to obtain target item information;
[0012] A product retrieval unit, configured to retrieve product information in a preset product information database according to the target item information, so as to obtain the product information of the target product;
[0013] A product push unit, configured to push the product information of the target product to the first object on the screen interface of the current video playback.
[0014] Optionally, the item recognition unit is further configured to:
[0015] Determine a target video frame corresponding to the current video playback content according to the target text information;
[0016] Perform item recognition on the target video frame according to the target text information to obtain target item information.
[0017] Optionally, the item recognition unit is further configured to:
[0018] Determine a first time interval where the target text information is located, and determine multiple candidate video frames corresponding to the current video playback content according to the first time interval;
[0019] Perform frame screening processing on the multiple candidate video frames to obtain a target video frame corresponding to the current video playback content.
[0020] Optionally, the item recognition unit is further configured to:
[0021] Perform frame deduplication and clarity screening on the multiple candidate video frames to obtain multiple first video frames meeting the clarity requirements;
[0022] Perform frame screening on the multiple first video frames for the target face to obtain multiple second video frames including the target face;
[0023] Perform frame screening on the multiple second video frames for the human body area range corresponding to the target face to obtain third video frames meeting the human body area range requirements;
[0024] Use the third video frame as the target video frame corresponding to the current video playback content.
[0025] Optionally, the item recognition unit is further configured to:
[0026] Perform frame clustering based on similarity on the multiple candidate video frames to obtain multiple video frame sets. In each video frame set, retain one of the candidate video frames and delete the other candidate video frames in the video frame set;
[0027] Alternatively, perform frame separation on the multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; in each of the video frame intervals, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval.
[0028] Alternatively, perform frame separation on the multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; calculate the similarity between every two adjacent candidate video frames in each of the video frame intervals; when all the similarities are greater than or equal to a similarity threshold, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval; when there is a case where any of the similarities is less than the similarity threshold, retain the two candidate video frames with similarities less than the similarity threshold and delete the other candidate video frames in the video frame interval.
[0029] Optionally, the screen interface includes a video playback interface area and a product push interface area. The video playback interface area displays the current video playback content, and the product push interface area displays the product information of the target product. The product push device further includes:
[0030] A first page jump unit, configured to, when detecting that the first object triggers the product information of the target product, keep displaying the current video playback content in the video playback interface area and jump to display the product purchase page of the target product in the product push interface area;
[0031] A first page processing unit, configured to, when detecting that the first object has completed the product purchase of the target product on the product purchase page, not display the product push interface area.
[0032] Optionally, the product push device further includes:
[0033] A second page jump unit, configured to, when detecting that the first object triggers the product information of the target product, pause the currently playing video and jump the screen interface to display the product purchase page of the target product;
[0034] A second page processing unit, configured to, when detecting that the first object has completed the product purchase of the target product on the product purchase page, jump the screen interface to display the original video playback page and continue playing the current video.
[0035] Optionally, the product push device further includes a product information library construction unit, and the product information library construction unit is configured to:
[0036] Obtain video samples, perform frame interval screening on the video samples to obtain multiple video sample frame intervals;
[0037] Perform frame screening processing on multiple of the video sample frame intervals respectively to obtain multiple target video sample frames;
[0038] Perform item recognition on each of the target video sample frames to obtain target item sample information in each of the target video sample frames;
[0039] Perform product information retrieval on each of the target item sample information to obtain product information of the product samples corresponding to each of the target item sample information;
[0040] Construct the product information library according to the product information of the video samples and the product samples.
[0041] Optionally, the product information library construction unit is further configured to:
[0042] Obtain multiple barrage information of the video sample, perform semantic recognition on each of the barrage information to obtain multiple descriptive words related to the items appearing in the video sample;
[0043] Determine the second time interval where each of the descriptive words is located, and perform frame interval screening on the video sample according to the multiple second time intervals to obtain multiple video sample frame intervals.
[0044] Optionally, the product information library construction unit is further configured to:
[0045] Perform frame deduplication and clarity screening on each of the video sample frame intervals to obtain multiple first video sample frames that meet the clarity requirements in each of the video sample frame intervals;
[0046] Perform frame screening on the multiple first video sample frames according to the preset object face to obtain multiple second video sample frames including the object face;
[0047] Perform frame screening on the multiple second video sample frames for the human body area range corresponding to the object face to obtain multiple target video sample frames that meet the human body area range requirements.
[0048] Optionally, the video sample frame interval includes multiple candidate video sample frames; the product information library construction unit is further configured to:
[0049] In each of the video sample frame intervals, perform frame clustering based on similarity to obtain multiple video sample frame sets, and in each of the video sample frame sets, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame set;
[0050] Alternatively, in each of the video sample frame intervals, perform frame interval segmentation according to a preset frame sequence difference to obtain a plurality of video sample frame sub-intervals; in each of the video sample frame sub-intervals, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame sub-interval;
[0051] Alternatively, in each of the video sample frame intervals, perform frame interval segmentation according to a preset frame sequence difference to obtain a plurality of video sample frame sub-intervals; in each of the video sample frame sub-intervals, calculate the similarity between every two adjacent candidate video sample frames; when all the similarities are greater than or equal to a similarity threshold, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame sub-interval; when there is a case where any of the similarities is less than the similarity threshold, retain the two candidate video sample frames with similarities less than the similarity threshold and delete the other candidate video sample frames in the video sample frame sub-interval.
[0052] On the other hand, an embodiment of the present application further provides an electronic device, including:
[0053] At least one processor;
[0054] At least one memory for storing at least one program;
[0055] When the at least one program is executed by the at least one processor, the commodity push method described above is implemented.
[0056] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored, and when the computer program executable by the processor is executed by the processor, it is used to implement the commodity push method described above.
[0057] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program or computer instructions, the computer program or the computer instructions are stored in a computer-readable storage medium, a processor of an electronic device reads the computer program or the computer instructions from the computer-readable storage medium, and the processor executes the computer program or the computer instructions, so that the electronic device executes the commodity push method described above.
[0058] The embodiments of the present application at least include the following beneficial effects: During the video playback, when the target text information input by the first object is received, item recognition is performed on the current video playback content according to the target text information to obtain the target item information. Since the target item information is obtained by performing item recognition on the current video playback content according to the target text information input by the first object, the target item information can meet the preferences and needs of the first object. Then, according to the target item information, a product information search is performed in the preset product information database to obtain the product information of the target product, and on the screen interface of the current video playback, the product information of the target product is pushed to the first object. Since the product information of the target product is obtained by performing a product information search in the preset product information database according to the target item information, the target product can better meet the purchase preferences and purchase needs of the first object. In addition, by pushing the product information of the target product to the first object on the screen interface of the current video playback, not only does it not require switching the current video playback interface to the shopping platform, but also it does not require performing a product search for the target product on the shopping platform. Therefore, the cost for the first object to obtain product information that meets their purchase preferences and purchase needs can be reduced without affecting the viewing integrity of the video content for the first object. Furthermore, since the target product pushed on the screen interface of the current video playback is retrieved according to the target text information input by the first object, the shopping needs of the first object when watching the video can be better met, enabling the first object to more easily obtain the product information of the corresponding product according to their purchase preferences or purchase needs, thereby effectively improving the shopping satisfaction of the first object.
[0059] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. They are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0061] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0062] Figure 2 It is a schematic diagram of another implementation environment provided by an embodiment of the present application;
[0063] Figure 3 It is a flowchart of a product pushing method provided by an embodiment of the present application;
[0064] Figure 4 It is a schematic diagram of the target text information input by a first object provided by an embodiment of the present application;
[0065] Figure 5 It is a schematic diagram of the target text information input by a first object provided by another embodiment of the present application;
[0066] Figure 6 It is a schematic diagram of the construction process of a commodity information library provided by a specific example of the present application;
[0067] Figure 7 It is a schematic diagram of the process of screening target video sample frames in a video sample provided by an embodiment of the present application;
[0068] Figure 8 It is a schematic diagram of the process of performing commodity information retrieval provided by an embodiment of the present application;
[0069] Figure 9 It is a schematic diagram of a screen interface when performing a commodity push method provided by an embodiment of the present application;
[0070] Figure 10 It is a schematic diagram of another screen interface when performing a commodity push method provided by an embodiment of the present application;
[0071] Figure 11 It is a schematic diagram of another screen interface when performing a commodity push method provided by an embodiment of the present application;
[0072] Figure 12 It is a detailed flowchart of a commodity push method provided by a specific example of the present application;
[0073] Figure 13 It is a schematic diagram of a commodity push device provided by an embodiment of the present application;
[0074] Figure 14 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0075] The present application will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0076] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0078] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0079] 1) Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0080] 2) Computer Vision Technology (CV): Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0081] 3) Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0082] 4) Video structured data refers to data extracted from videos that contains structured information. Video structured data can include information such as objects, faces, scenes, actions, and voices in the video, and this information can be used for applications such as video content analysis, video retrieval, video classification, video summarization, and video annotation. Video structured data is usually extracted from videos through computer vision and machine learning technologies and then processed and analyzed.
[0083] 5) Multi-label Classification is a classification task where each sample may have multiple labels. Different from traditional single-label classification, the challenge in multi-label classification is that a sample may belong to multiple categories simultaneously, and there may be dependencies between different categories, which makes the multi-label classification problem more complex and challenging.
[0084] 6) Multimodal retrieval is an information retrieval technology that aims to process and integrate information from different modalities (such as text, images, audio, videos, etc.) to achieve more accurate, efficient, and comprehensive retrieval results. Multimodal retrieval has a wide range of applications, such as being able to be applied in smart homes, multimedia databases, intelligent search, and other fields.
[0085] Currently, when users are watching videos and have a purchasing intention for the products shown in the videos, for example, when they hope to buy the same clothes worn by the protagonist in the video, they will use a shopping platform integrated with product retrieval and recommendation functions to search for and purchase the same clothes. Although this shopping platform integrates product retrieval and recommendation functions and can provide consumers with a convenient shopping experience, the product retrieval and recommendation functions provided by the shopping platform have no direct connection with the video content itself. Users cannot obtain relevant product information during the process of watching the video and need to jump to the shopping platform from the current video playback interface to search for relevant product information. Therefore, throughout the process, it will not only affect the integrity of users' viewing of video content but also increase the cost for users to obtain product information.
[0086] In order not to affect the viewing integrity of the video content for users and reduce the cost for users to obtain product information, a solution of combining product information with video content has been proposed in the related art, enabling users to obtain relevant product information during the video viewing process. For example, product information of relevant products is associated with a certain frame of the video, and when the video frame is played, the corresponding product information is displayed, or a series of product information of products related to the video content is presented after the video ends. However, this product push method depends on the subjective will and choice of the video publisher or merchant, which easily leads to the product information displayed during the video playback may be product information that does not meet the interests and needs of users. That is to say, users cannot obtain the product information of corresponding products according to their own preferences or will during the video playback. For example, there are two products, clothes and trousers, in the video playback content, and the user is interested in the trousers, but the product information of the clothes is displayed during the video playback. In this case, if the user wants to obtain the product information of the trousers, they still need to jump to the shopping platform from the current video playback interface for product retrieval to find the corresponding product information of the trousers. Therefore, there will be a problem of inaccurate product push in this related technology, which will increase the cost for users to obtain product information.
[0087] In order to enable users to obtain product information of corresponding products according to their purchase preferences or purchase needs without affecting the integrity of the user's viewing of video content, thereby reducing the cost for users to obtain product information that meets their purchase preferences and purchase needs, the embodiments of the present application provide a product push method, a product push device, an electronic device, a computer-readable storage medium, and a computer program product. During the video playback process, when receiving target text information input by the user, perform item recognition on the current video playback content according to the target text information to obtain target item information. Since the target item information is obtained by performing item recognition on the current video playback content according to the target text information input by the user, the target item information can conform to the user's preferences and needs. Then, retrieve product information in a pre-set product information library according to the target item information to obtain the product information of the target product, and push the product information of the target product to the user on the screen interface of the current video playback. Since the product information of the target product is obtained by retrieving product information in the pre-set product information library according to the target item information, the target product can better conform to the user's purchase preferences and purchase needs. In addition, by pushing the product information of the target product to the user on the screen interface of the current video playback, not only does it not require switching the current video playback interface to a shopping platform, but also it does not require retrieving the product in the shopping platform. Therefore, it can reduce the cost for users to obtain product information that meets their purchase preferences and purchase needs without affecting the integrity of the user's viewing of video content. Furthermore, since the target product pushed on the screen interface of the current video playback is retrieved according to the target text information input by the user, it can better meet the shopping needs of users when watching videos, enabling users to more easily obtain product information of corresponding products according to their purchase preferences or purchase needs, thereby effectively improving the shopping satisfaction of users.
[0088] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application. Refer to Figure 1 , this implementation environment includes a first user terminal 101 and a first server 102. The first user terminal 101 and the first server 102 are directly or indirectly connected through a wired or wireless communication method. Among them, the first user terminal 101 and the first server 102 can be nodes in a blockchain, and this embodiment does not make specific limitations on this.
[0089] The first user terminal 101 may include, but is not limited to, intelligent devices such as smart phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, and aircraft. Optionally, the first user terminal 101 may be installed with a video playback client. Through the video playback client, online videos or offline videos can be watched, and product information that meets the user's purchase preferences and purchase needs can be obtained during the video viewing process.
[0090] The first server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0091] In one embodiment, the first server 102 has at least functions such as text acquisition, item recognition, product retrieval, and product push. For example, the first server 102 can receive the target text information input by the user, perform item recognition on the current video playback content according to the target text information to obtain the target item information, then perform product information retrieval in the pre-set product information library according to the target item information to obtain the product information of the target product, and then push the product information of the target product to the user on the screen interface of the current video playback.
[0092] Refer to Figure 1 As shown, in an application scenario, assume that the first user terminal 101 is a smart phone, and the first user terminal 101 is installed with a video playback client. During the process of the user watching a video through the video playback client in the first user terminal 101, in response to the user inputting the target text information in the screen interface of the current video playback, the first user terminal 101 sends the target text information to the first server 102; in response to receiving the target text information, the first server 102 performs item recognition on the current video playback content according to the target text information to obtain the target item information, then performs product information retrieval in the locally pre-set product information library according to the target item information to obtain the product information of the target product, and then sends the product information of the target product along with the video frames of the currently played video to the first user terminal 101, so that the first user terminal 101 displays the product information of the target product in the screen interface of the current video playback while playing the video, achieving the purpose of pushing the product information of the target product to the user during the video viewing process.
[0093] Figure 2 It is a schematic diagram of another implementation environment provided by the embodiments of the present application. Refer to Figure 2, the implementation environment includes a second user terminal 201, a second server 202, and a database server 203. The second user terminal 201 and the second server 202 are directly or indirectly connected by wired or wireless communication means, and the second server 202 and the database server 203 are directly or indirectly connected by wired or wireless communication means. Among them, the second user terminal 201, the second server 202, and the database server 203 can all be nodes in the blockchain, and this embodiment does not make specific limitations on this.
[0094] The second user terminal 201 may include, but is not limited to, intelligent devices such as smart phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, and aircraft. Optionally, the second user terminal 201 may be installed with a video playback client, through which online videos or offline videos can be watched, and product information that meets the user's purchase preferences and purchase needs can be obtained during the video viewing process.
[0095] The second server 202 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN networks, and big data and artificial intelligence platforms. Among them, a product information library may be pre-deployed in the database server 203, and product information of multiple different products is pre-stored in the product information library. The second server 202 can obtain the product information of the corresponding target product from the product information library in the database server 203 according to the identified target item information.
[0096] In one embodiment, the second server 202 has at least functions such as text acquisition, item recognition, product retrieval, and product push. For example, the second server 202 can receive the target text information input by the user, perform item recognition on the current video playback content according to the target text information to obtain the target item information, and then send the target item information to the database server 203, so that the database server 203 performs product information retrieval in the product information library according to the target item information to obtain the product information of the target product, and then receives the product information of the target product sent by the database server 203, and pushes the product information of the target product to the user on the screen interface of the current video playback.
[0097] Refer to Figure 2As shown, in another application scenario, it is assumed that the second user terminal 201 is a computer, and a video playback client is installed on the second user terminal 201. During the process of the user watching a video through the video playback client in the second user terminal 201, in response to the user inputting target text information in the screen interface of the current video playback, the second user terminal 201 sends the target text information to the second server 202; in response to receiving the target text information, the second server 202 performs item recognition on the current video playback content according to the target text information to obtain target item information, and then sends the target item information to the database server 203; in response to receiving the target item information, the database server 203 calls the product information library corresponding to the currently played video, and retrieves product information in the product information library according to the target item information to obtain the product information of the target product, and then sends the product information of the target product to the second server 202; in response to receiving the product information of the target product, the second server 202 sends the product information of the target product along with the video frames of the current played video to the second user terminal 201, so that the second user terminal 201 displays the product information of the target product in the screen interface of the current video playback while playing the video, achieving the purpose of pushing the product information of the target product to the user during the process of the user watching the video.
[0098] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing according to the attribute information or set of attribute information of the target object (such as a user, etc.) and data related to the characteristics of the target object, the permission or consent of the target object will be obtained first, and moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the relevant data of the target object necessary for the normal operation of the embodiments of the present application will be obtained.
[0099] Figure 3 It is a flowchart of a product push method provided by an embodiment of the present application. This product push method can be executed by a server or a terminal, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the method being executed by the server as an example for illustration. Refer to Figure 3 and this product push method includes but is not limited to steps 310 to 330.
[0100] Step 310: During the video playback, receive the target text information input by the first object, and perform item recognition on the current video playback content according to the target text information to obtain target item information.
[0101] In one embodiment, the first object may be a user watching a video. The target text information input by the first object may be bullet screen information related to an item that appears in the video playback content and posted by the user, or may be item description information related to an item that appears in the video playback content and input by the user in the search bar, which is not specifically limited herein. For example Figure 4 As shown, assume that the first object (such as a user watching a video) posts a first bullet screen information 410 (such as Figure 4 "The heroine is really good-looking") for the current video playback content, a second bullet screen information 420 (such as Figure 4 "The skirt is really beautiful"), a third bullet screen information 430 (such as Figure 4 "The environment is really beautiful"), etc. Since the second bullet screen information 420 (such as Figure 4 "The skirt is really beautiful") among them is related to an item that appears in the current video playback content, it can be considered that the second bullet screen information 420 (such as Figure 4 "The skirt is really beautiful") is the target text information input by the first object. Another example is Figure 5 As shown, assume that during the process of watching a video, the user inputs description information of a commodity to be purchased through the search bar 510 on the screen interface, then it can be considered that the description information of the commodity is the target text information input by the first object.
[0102] In one embodiment, when performing item recognition on the current video playback content according to the target text information to obtain target item information, the target video frame corresponding to the current video playback content may be determined first according to the target text information, and then item recognition is performed on the target video frame according to the target text information to obtain the target item information. Since the target item information is obtained by performing item recognition on the target video frame corresponding to the current video playback content according to the target text information input by the first object, the target item information can meet the preferences and needs of the first object. Then when the subsequent steps perform commodity information retrieval according to the target item information, the commodity information of the target commodity that meets the purchase preferences and purchase needs of the first object can be retrieved more accurately.
[0103] In one embodiment, the target item information may include information such as the image, type, color, style of the target item, which is not specifically limited herein.
[0104] In one embodiment, since the target text information input by the first object will be maintained on the screen interface of the current video playback for a certain period of time. For example, the bullet screen information will move from the right side to the left side of the video playback interface. Therefore, the bullet screen information will be maintained on the screen interface of the current video playback for the time from moving from the right side to the left side of the video playback interface. During this time, multiple video frames will be played. Therefore, in the process of determining the target video frame corresponding to the current video playback content according to the target text information, the first time interval where the target text information is located can be determined first, multiple candidate video frames corresponding to the current video playback content can be determined according to the first time interval, and then frame screening processing is performed on the multiple candidate video frames to obtain the target video frame corresponding to the current video playback content. Among them, the first time interval where the target text information is located can be the retention time of the target text information on the screen interface of the current video playback. Screening the target video frame corresponding to the current video playback content through the first time interval where the target text information is located can effectively improve the accuracy of obtaining the target video frame corresponding to the current video playback content.
[0105] In one embodiment, in the process of performing frame screening processing on multiple candidate video frames to obtain the target video frame corresponding to the current video playback content, frame deduplication and clarity screening can be performed on the multiple candidate video frames first to obtain multiple first video frames that meet the clarity requirements, then frame screening for the target face is performed on the multiple first video frames to obtain multiple second video frames including the target face, and then frame screening for the human body area range corresponding to the target face is performed on the multiple second video frames to obtain third video frames that meet the requirements of the human body area range. Then, the third video frames are used as the target video frames corresponding to the current video playback content. Among them, the target face can be a pre-collected star face or the face of the main character in the video, such as the faces of the male and female protagonists or the male and female supporting roles in the video, etc., which is not specifically limited here. When performing frame screening for the target face on the multiple first video frames, a face library including the face information of multiple target faces such as star faces or the faces of the main characters in the video can be pre-constructed, and then a deep learning algorithm or a traditional feature extraction algorithm is used to identify the faces in each of the first video frames to obtain the face recognition results. Then, the face recognition results of each of the first video frames are compared with the face information in the pre-constructed face library to determine whether the target faces in the face library appear in these first video frames. Then, the first video frames in which the target faces in the face library appear are screened out, and multiple second video frames including the target face can be obtained.
[0106] In one embodiment, after obtaining a plurality of second video frames each including a target human face, although all of these second video frames include the target human face, there may be a situation where some of these second video frames only include the target human face but do not include the merchandise that the first object is interested in. Therefore, after obtaining a plurality of second video frames each including a target human face, it is also necessary to perform frame screening on these second video frames for the human body region range corresponding to the target human face to obtain third video frames that meet the requirements of the human body region range. By performing frame screening on these second video frames for the human body region range corresponding to the target human face, third video frames with complete clothing (i.e., meeting the requirements of the human body region range) can be obtained, which is beneficial for subsequent using the third video frames as the target video frames corresponding to the current video playback content for item recognition. Among them, meeting the requirements of the human body region range may mean that the human body range appearing in the video frame meets a preset human body region range. The preset human body region range may include a head range, an upper body range, a lower body range, or an entire human body range, etc., and can be appropriately selected according to the actual application situation, and no specific limitation is made here. Therefore, in the process of performing frame screening on a plurality of second video frames for the human body region range corresponding to the target human face to obtain third video frames that meet the requirements of the human body region range, human body detection can be performed on the people in each second video frame to obtain the human body contour points and human body key points (such as head key points, neck key points, waist key points, hip key points, etc.) of the people in each second video frame, and then determine whether the people in the second video frame are wearing complete clothing based on the obtained human body contour points and human body key points. For example, assuming that the obtained human body contour points include head contour points and the obtained human body key points include head key points, then it can be considered that the person in the second video frame is wearing a complete hat; another example, assuming that the obtained human body contour points include upper body contour points and the obtained human body key points include waist key points, then it can be considered that the person in the second video frame is wearing complete clothes; yet another example, assuming that the obtained human body contour points include the entire human body contour points and the obtained human body key points include foot key points, then it can be considered that the person in the second video frame is wearing complete clothes and pants. After obtaining the third video frames that meet the requirements of the human body region range, the third video frames can be used as the target video frames corresponding to the current video playback content, and thus item recognition can be performed on the target video frames according to the target text information to obtain the target item information in the target video frames.
[0107] In one embodiment, when performing item recognition on a target video frame according to target text information to obtain target item information, semantic recognition can be first performed on the target text information to obtain target text semantic information, and then item recognition is performed on the target video frame to obtain multiple candidate item information in the target video frame. Next, among the multiple candidate item information, the target item information that matches the target text semantic information is determined. Among them, there may be multiple candidate item information in the target video frame. For example, it may include information about items such as hats, clothes, pants, shoes, etc. However, for the user (i.e., the first object), what they are more concerned about is often one of these candidate item information. Since the target text information input by the first object is the description information of the item they are more concerned about, the target item information with a higher matching degree can be determined among the multiple candidate item information according to the target text semantic information of the target text information. Then, when the subsequent steps perform commodity information retrieval according to the target item information, the commodity information of the target commodity that meets the purchase preferences and purchase needs of the first object can be retrieved more accurately.
[0108] In one embodiment, in the process of frame deduplication for multiple candidate video frames, frame clustering based on similarity can be first performed on the multiple candidate video frames to obtain multiple video frame sets, and then in each video frame set, one of the candidate video frames is retained, and other candidate video frames in the video frame set are deleted. Among them, when performing frame clustering based on similarity on the multiple candidate video frames to obtain multiple video frame sets, the similarity between each candidate video frame and all other candidate video frames can be calculated, and then multiple candidate video frames with a similarity greater than or equal to the similarity threshold are subjected to frame clustering to obtain multiple video frame sets. Since the candidate video frames in each video frame set are similar, in order to reduce the processing pressure of subsequent item recognition and improve the efficiency of item recognition, only one candidate video frame can be retained in each video frame set. Therefore, for each video frame set, one of the candidate video frames can be retained, and then other candidate video frames in the video frame set are deleted. In one embodiment, the similarity threshold can be appropriately selected according to the actual application situation, and no specific limitation is made here. For example, the similarity threshold can be set to 0.9 or 0.95.
[0109] In one embodiment, during the process of frame deduplication for multiple candidate video frames, the multiple candidate video frames may first be frame-separated according to a preset frame sequence difference to obtain multiple video frame intervals, and then in each video frame interval, one of the candidate video frames is retained, and the other candidate video frames in the video frame interval are deleted. Among multiple candidate video frames within the preset frame sequence difference, it can be considered that the change in the frame images between them is very small. That is to say, the multiple candidate video frames in each video frame interval obtained by frame-separating the multiple candidate video frames according to the preset frame sequence difference can be considered to be similar. In order to reduce the processing pressure of subsequent item recognition and improve the efficiency of item recognition, only one candidate video frame can be retained in each video frame interval. Therefore, for each video frame interval, one of the candidate video frames can be retained, and then the other candidate video frames in the video frame interval are deleted. In one embodiment, the preset frame sequence difference can be appropriately selected according to the actual application situation, and no specific limitation is provided here. For example, the preset frame sequence difference can be set to 50 or 45.
[0110] In one embodiment, during the process of frame deduplication for multiple candidate video frames, the multiple candidate video frames can first be frame-separated according to a preset frame sequence difference to obtain multiple video frame intervals. Then, in each video frame interval, the similarity between every two adjacent candidate video frames is calculated. If all similarities are greater than or equal to the similarity threshold, one of the candidate video frames is retained, and the other candidate video frames in the video frame interval are deleted. If there is a case where the similarity is less than the similarity threshold among all similarities, the two candidate video frames with similarities less than the similarity threshold are retained, and the other candidate video frames in the video frame interval are deleted. Among multiple candidate video frames within the preset frame sequence difference, although in most cases it can be considered that the change in the frame images among them is very small, for the case where there is a video frame with a shot transition among multiple candidate video frames within the preset frame sequence difference, it cannot be considered that the change in the frame images among these candidate video frames is very small. Obviously, the similarity between the two candidate video frames before and after the shot transition is very small. To avoid errors caused by this situation, after the multiple candidate video frames are frame-separated according to the preset frame sequence difference to obtain multiple video frame intervals, for each video frame interval, the similarity between every two adjacent candidate video frames is first calculated. If all similarities are greater than or equal to the similarity threshold, it indicates that there is no video frame with a shot transition in this video frame interval. Therefore, one of the candidate video frames in this video frame interval can be retained, and the other candidate video frames in this video frame interval are deleted. If there is a case where the similarity is less than the similarity threshold among all similarities, it indicates that there is a video frame with a shot transition in this video frame interval. Therefore, the two candidate video frames with similarities less than the similarity threshold in this video frame interval can be retained, and the other candidate video frames in this video frame interval are deleted. In one embodiment, both the preset frame sequence difference and the similarity threshold can be appropriately selected according to the actual application situation, and no specific limitation is made here. For example, the preset frame sequence difference can be set to 50 or 45, and the similarity threshold can be set to 0.9 or 0.95.
[0111] Step 320: Retrieve product information in the preset product information database according to the target item information to obtain the product information of the target product.
[0112] In one embodiment, after obtaining the target item information in the current video playback content, the product information of the target product can be retrieved in the preset product information database according to the target item information. Since the product information of the target product is obtained by retrieving product information in the preset product information database according to the target item information, the target product can better meet the purchase preferences and purchase needs of the first object.
[0113] In one embodiment, the product information of the target product may include the image information, label information, product link, etc. of the target product. The label information of the target product may include information such as type, color, style, etc., which is not specifically limited herein.
[0114] In one embodiment, the product information of multiple candidate products may be pre - stored in the product information database. When the target item information is descriptive information such as type, color, style, etc., and the product information of the candidate products is label information, when performing product information retrieval in the pre - set product information database according to the target item information, the target item information can be compared with the product information of multiple candidate products in the product information database, and among the product information of multiple candidate products in the product information database, the product information of the target product that matches the target item information can be found. Additionally, when both the target item information and the product information of the candidate products are image information, when performing product information retrieval in the pre - set product information database according to the target item information, the image similarity between the target item information and the product information of multiple candidate products in the product information database can be calculated, and then the product information of the candidate product with the highest image similarity can be used as the product information of the target product.
[0115] In one embodiment, the product information database is constructed according to the following steps: First, obtain video samples, perform frame interval screening on the video samples to obtain multiple video sample frame intervals, then perform frame screening processing on each of the multiple video sample frame intervals to obtain multiple target video sample frames, and perform item recognition on each target video sample frame to obtain the target item sample information in each target video sample frame. Then, perform product information retrieval on each target item sample information to obtain the product information of the product samples corresponding to each target item sample information. Next, construct the product information database according to the video samples and the product information of the product samples. Among them, when constructing the product information database according to the video samples and the product information of the product samples, the product information database can be constructed according to the video ID of the video samples and the product information of the product samples.
[0116] In one embodiment, the video sample can be a complete long video or short video, or a video clip, which is not specifically limited herein. Additionally, the target item sample information may include information such as the image, type, color, style, etc. of the target item sample, which is not specifically limited herein. Furthermore, the product information of the product sample may include image information, label information, product link, etc. of the product sample. The label information of the product sample may include information such as type, color, style, etc., which is not specifically limited herein.
[0117] In one embodiment, when screening frame intervals of a video sample to obtain multiple video sample frame intervals, multiple barrage messages of the video sample can be obtained first, and semantic recognition can be performed on each barrage message to obtain multiple descriptive words related to the items appearing in the video sample. Then, the second time interval where each descriptive word is located can be determined, and the video sample can be screened for frame intervals according to the multiple second time intervals to obtain multiple video sample frame intervals. For example, multiple barrage messages of the video sample can be obtained first, and then semantic recognition can be performed on these barrage messages respectively to extract descriptive words related to the items appearing in the video sample, such as "clothes", "coat", "nice", "beautiful", etc. Then, according to the second time intervals where these descriptive words are located, the video sample can be screened for frame intervals to obtain video sample frame intervals corresponding to these descriptive words respectively. Among them, the second time interval where the descriptive word is located can be the retention time of the barrage message corresponding to the descriptive word on the screen interface during video playback. These descriptive words in the barrage message can reflect the real-time feedback and needs of the audience for the video content. Therefore, by extracting these descriptive words from the barrage message and analyzing the time points when these descriptive words appear in the video, not only can the focus of the audience's interest be found more accurately, but also the interference of redundant pictures in the video sample can be reduced, which is beneficial to improving the accuracy of commodity push.
[0118] In one embodiment, when there is no barrage message or the number of barrage messages is small in the video sample, when screening the video sample for frame intervals to obtain multiple video sample frame intervals, the method of evenly extracting frames can be used to screen and obtain multiple video sample frame intervals. For example, every first number of video sample frames can be skipped in the video sample to extract the second number of video sample frames, and the second number of video sample frames extracted each time can be used as a video sample frame interval, so that multiple video sample frame intervals can be obtained. Among them, both the first number and the second number can be appropriately selected according to the actual application situation, and no specific limitation is made here. For example, the first number can be 500 and the second number can be 300.
[0119] In one embodiment, in the process of performing frame screening on multiple video sample frame intervals respectively to obtain multiple target video sample frames, frame deduplication and clarity screening can be first performed on each video sample frame interval to obtain multiple first video sample frames that meet the clarity requirements in each video sample frame interval. Then, frame screening is performed on the multiple first video sample frames according to the preset object face to obtain multiple second video sample frames including the object face. Next, frame screening is performed on the multiple second video sample frames for the human body area range corresponding to the object face to obtain multiple target video sample frames that meet the human body area range requirements. Among them, the object face can be a pre-collected star face or the face of the main character in the video sample, such as the faces of the male and female protagonists or the male and female supporting roles in the video sample, etc., which is not specifically limited here. When performing frame screening on the multiple first video sample frames according to the preset object face, a face library including the face information of multiple object faces such as star faces or the faces of the main characters in the video sample can be pre-constructed. Then, a deep learning algorithm or a traditional feature extraction algorithm is used to identify the faces in each first video sample frame to obtain a face recognition result. Then, the face recognition results of each first video sample frame are compared with the face information in the pre-constructed face library to determine whether the object faces in the face library appear in these first video sample frames. Then, the first video sample frames with the object faces in the face library are screened out, and multiple second video sample frames including the object face can be obtained.
[0120] In one embodiment, since a face library including the face information of multiple object faces such as star faces or the faces of the main characters in the video sample can be pre-constructed in the process of obtaining the second video sample frames, in this case, when constructing a product information library according to the product information of the video sample and the product sample, the product information library can be constructed according to the product information of the video sample, the product sample, and the face information of the object face in this face library. That is to say, the constructed product information library can include the product information of the video sample, the product sample, and the face information of the object face.
[0121] In one embodiment, after obtaining a plurality of second video sample frames including the face of the object, although these second video sample frames all include the face of the object, there may be a situation in these second video sample frames that only include the face of the object but not the commodity. Therefore, after obtaining a plurality of second video sample frames including the face of the object, it is also necessary to perform frame screening on these second video sample frames for the human body region range corresponding to the face of the object, and obtain a plurality of target video sample frames that meet the requirements of the human body region range. By performing frame screening on these second video sample frames for the human body region range corresponding to the face of the object, a target video sample frame with complete clothing (i.e., meeting the requirements of the human body region range) can be obtained, which is conducive to the subsequent object recognition of the target video sample frame. Among them, meeting the requirements of the human body region range may mean that the human body range appearing in the video sample frame meets the preset human body region range, wherein the preset human body region range may include the head range, the upper body range, the lower body range or the entire human body range, etc., which can be appropriately selected according to the actual application situation, and is not specifically limited here. Therefore, in the process of performing frame screening on multiple second video sample frames for the human body region range corresponding to the target face to obtain multiple target video sample frames that meet the human body region range requirements, human body detection can be performed on the characters in each second video sample frame to obtain the human body contour points and human body key points (such as head key points, neck key points, waist key points, hip key points, etc.) of the characters in each second video sample frame, and then determine whether the characters in the second video sample frame are fully dressed based on the obtained human body contour points and human body key points. For example, assuming that the obtained human body contour points include head contour points and the obtained human body key points include head key points, then it can be considered that the characters in the second video sample frame are wearing a complete hat; for another example, assuming that the obtained human body contour points include upper body contour points and the obtained human body key points include waist key points, then it can be considered that the characters in the second video sample frame are wearing complete clothes; for another example, assuming that the obtained human body contour points include the entire human body contour points and the obtained human body key points include foot key points, then it can be considered that the characters in the second video sample frame are wearing complete clothes and pants. After obtaining a plurality of target video sample frames that meet the requirements of the human body region range, object recognition can be performed on each target video sample frame to obtain target object sample information in each target video sample frame.
[0122] In one embodiment, when performing object recognition on each target video sample frame to obtain the target object sample information in each target video sample frame, object recognition can be performed on each target video sample frame separately to obtain the types and location information of all object samples (such as hats, clothes, pants, shoes, etc.) in each target video sample frame. Then, for each target video sample frame, the occurrence range of each object sample in this target video sample frame can be obtained according to the location information, and the object sample with the largest occurrence range is used as the target object sample in this target video sample frame. Next, label information of this target object sample can be generated. For example, a multi-label classification model is called to generate label information for this target object sample, where the label information can include information such as color and style. At this time, the type, color, style, etc. information of this target object sample can be used as the target object sample information of this target object sample, that is, the target object sample information in each target video sample frame can be obtained.
[0123] In one embodiment, after obtaining the target object sample information in each target video sample frame, the commodity information of the commodity samples corresponding to each target object sample information can be retrieved, so as to obtain the commodity information of the commodity samples corresponding to each target object sample information. Among them, when retrieving the commodity information of each target object sample information, for each target object sample information, first, according to the location information of the target object sample corresponding to this target object sample information, the image information of the corresponding target object sample can be obtained in the corresponding target video sample frame, and then a multi-modal retrieval model is called to perform commodity information retrieval according to this target object sample information and this image information to obtain the commodity information of the commodity sample most similar to this target object sample information. In one embodiment, since the label information of the target object samples in each target video sample frame can be obtained when performing object recognition on each target video sample frame, in this case, when constructing a commodity information library according to the commodity information of video samples and commodity samples, in addition to constructing a commodity information library according to the commodity information of video samples, commodity samples and the face information of object faces in the face library, a commodity information library can also be constructed according to the commodity information of video samples, commodity samples, the face information of object faces in the face library and the label information of target object samples, that is, the constructed commodity information library can include the commodity information of video samples, commodity samples, the face information of object faces and the label information of target object samples.
[0124] In one embodiment, a video sample frame interval may include multiple candidate video sample frames. In this case, when performing frame deduplication on each video sample frame interval, in each video sample frame interval, first, frame clustering based on similarity can be performed to obtain multiple video sample frame sets. Then, in each video sample frame set, one candidate video sample frame can be retained, and other candidate video sample frames in the video sample frame set can be deleted. Among them, when performing frame clustering based on similarity in the video sample frame interval to obtain multiple video sample frame sets, the similarity between each candidate video sample frame in the video sample frame interval and all other candidate video sample frames can be calculated. Then, multiple candidate video sample frames with a similarity greater than or equal to the similarity threshold are subjected to frame clustering to obtain multiple video sample frame sets. Since the candidate video sample frames in each video sample frame set are similar, in order to reduce the processing pressure of subsequent item recognition and improve the efficiency of item recognition, only one candidate video sample frame can be retained in each video sample frame set. Therefore, for each video sample frame set, one candidate video sample frame can be retained, and other candidate video sample frames in the video sample frame set can be deleted. In one embodiment, the similarity threshold can be appropriately selected according to the actual application situation, and no specific limitation is provided here. For example, the similarity threshold can be set to 0.9 or 0.95.
[0125] In one embodiment, a video sample frame interval may include multiple candidate video sample frames. In this case, when performing frame deduplication on each video sample frame interval, in each video sample frame interval, the video sample frame interval can also be segmented into multiple video sample frame sub-intervals according to a preset frame sequence difference first. Then, in each video sample frame sub-interval, one candidate video sample frame can be retained, and other candidate video sample frames in the video sample frame sub-interval can be deleted. Among them, for multiple candidate video sample frames within the preset frame sequence difference, it can be considered that the change in the frame images between them is very small. That is to say, for multiple candidate video sample frames in each video sample frame sub-interval obtained by segmenting the video sample frame interval according to the preset frame sequence difference, they can be considered similar. In order to reduce the processing pressure of subsequent item recognition and improve the efficiency of item recognition, only one candidate video sample frame can be retained in each video sample frame sub-interval. Therefore, for each video sample frame sub-interval, one candidate video sample frame can be retained, and other candidate video sample frames in the video sample frame sub-interval can be deleted. In one embodiment, the preset frame sequence difference can be appropriately selected according to the actual application situation, and no specific limitation is provided here. For example, the preset frame sequence difference can be set to 50 or 45.
[0126] In one embodiment, a video sample frame interval may include multiple candidate video sample frames. In this case, when performing frame deduplication on each video sample frame interval, in each video sample frame interval, multiple video sample frame sub-intervals may first be obtained by segmenting the frame interval according to a preset frame sequence difference. Then, in each video sample frame sub-interval, the similarity between every two adjacent candidate video sample frames is calculated. If all similarities are greater than or equal to the similarity threshold, one of the candidate video sample frames is retained, and other candidate video sample frames in the video sample frame sub-interval are deleted; if there is a case where a similarity is less than the similarity threshold among all similarities, two candidate video sample frames with similarities less than the similarity threshold are retained, and other candidate video sample frames in the video sample frame sub-interval are deleted. Among multiple candidate video sample frames within the preset frame sequence difference, although in most cases, it can be considered that the change in the frame images among them is very small, for the case where there is a video sample frame with a shot transition among multiple candidate video sample frames within the preset frame sequence difference, it cannot be considered that the change in the frame images among these candidate video sample frames is very small. Obviously, the similarity between the two candidate video sample frames before and after the shot transition is very small. To avoid errors caused by this situation, after segmenting the video sample frame interval into multiple video sample frame sub-intervals according to the preset frame sequence difference, for each video sample frame sub-interval, the similarity between every two adjacent candidate video sample frames is first calculated. If all similarities are greater than or equal to the similarity threshold, it indicates that there is no video sample frame with a shot transition in this video sample frame sub-interval. Therefore, one candidate video sample frame in this video sample frame sub-interval can be retained, and other candidate video sample frames in this video sample frame sub-interval are deleted; if there is a case where a similarity is less than the similarity threshold among all similarities, it indicates that there is a video sample frame with a shot transition in this video sample frame sub-interval. Therefore, two candidate video sample frames with similarities less than the similarity threshold in this video sample frame sub-interval can be retained, and other candidate video sample frames in this video sample frame sub-interval are deleted. In one embodiment, both the preset frame sequence difference and the similarity threshold can be appropriately selected according to the actual application situation, and no specific limitation is made here. For example, the preset frame sequence difference can be set to 50 or 45, and the similarity threshold can be set to 0.9 or 0.95.
[0127] The following uses a specific example to illustrate the construction process of the product information database in detail.
[0128] Refer to Figure 6 as shown Figure 6 is a schematic diagram of the construction process of the product information database provided by a specific example of this application. In Figure 6Among them, the construction process of the product information library can include three parts. The first part 610 is the preliminary preparation, the second part 620 is the ability dependence, that is, the deep learning ability required to establish the product information library, and the third part 630 is the information storage. After the information storage in the third part is completed, the required product information library can be established.
[0129] In the first part 610, the preliminary preparation can include the preparation of the face library, the preparation of the product library, and the preparation of the video samples. Among them, the preparation of the face library means that it is necessary to prepare the face information of the stars or the face information of the main characters in the video samples. For example, the face information of the male and female protagonists or the male and female supporting roles in the video samples. Specifically, the images of each star or the main characters in the video samples can be collected, and then these images can be screened and processed for face information extraction to obtain the face information of each star or the main characters in the video samples. Then, these face information can be classified and labeled, and the classified and labeled face information can be stored in the face library to complete the preparation of the face library. The preparation of the product library means that it is necessary to prepare the product information of various products, such as the types, colors, styles, images, product links, etc. of various products. Specifically, the product information such as the pictures, product descriptions, and product prices of each product can be collected, and then these product information can be classified and labeled. Then, the classified and labeled product information can be stored in the product library to complete the preparation of the product library. The preparation of the video samples means that it is necessary to prepare each video sample. Among them, these video samples can be complete long videos or short videos, or video clips, and no specific limitation is made here. In addition, when the video sample is a complete long video, in order to facilitate subsequent processing of the video sample, the complete long video can be sliced to obtain multiple video sample segments.
[0130] In the second part 620, the deep learning capabilities required to establish a product information database may include face recognition, human detection, object detection, and product retrieval, etc. Among them, face recognition is used to recognize the faces in the video sample frames of video samples. For example, deep learning algorithms or traditional feature extraction algorithms can be used to recognize the faces in the video sample frames to obtain face recognition results, and then the face recognition results are compared with the face information in the constructed face database to determine whether the face of a star or the main character appears in the video sample frame. For the video sample frames in which the face of a star or the main character is recognized, human detection can be further performed to determine whether the clothing worn by the star or the main character in the video sample frame is complete. When performing human detection, the human contour points and human key points (such as head key points, neck key points, waist key points, hip key points, etc.) of the star or the main character in the video sample frame can be detected, so that it can be judged whether the clothing worn by the star or the main character in the video sample frame is complete according to the detected human contour points and human key points. In addition, when performing human detection, a corresponding human detection box can be obtained. This human detection box can not only be used to obtain human contour points and human key points, but also provide an image cropping range for subsequent product retrieval. After completing human detection and obtaining video sample frames with complete clothing worn by the characters, object detection processing can be further performed. By performing object detection on the video sample frames with complete clothing worn by the characters, the position areas of items such as tops, pants, hats, etc. in the video sample frames and the label information of these items (such as label information including types, colors, styles, etc.) can be detected, so as to provide data preparation for subsequent fine-grained product retrieval. After completing object detection, product retrieval with double matching for the images and descriptions of each detected item can be performed to obtain the product information (such as including product links) of product samples that match both the image and the description.
[0131] In the third part 630, according to the ID of the video sample, the ID of the star face, or the ID of the main character face in the video sample obtained in the previous first part 610, and the label information and the product information (such as including product links) of the product samples obtained in the previous second part 620, a product information database corresponding to the video sample can be constructed. Among them, the label information in the product information database can be used to splice and generate the product push title content corresponding to the product. For example, it can be used to generate product push title content such as "XX same style blue suit", "XX same style floral long dress", etc.
[0132] In one embodiment, when using face recognition and human body detection in the second part 620 to screen target video sample frames in a video sample, the target video sample frames can be screened by extracting relevant structured information, where the structured information may include bullet screen information, face recognition result information, and human body detection result information of the video sample. Refer to Figure 7 As shown Figure 7 An exemplary flowchart for screening target video sample frames in a video sample is given. In Figure 7 it, the process of screening target video sample frames in a video sample mainly includes three processing procedures: video frame extraction 710, clustering and duplicate removal 720, and basic filtering 730.
[0133] In the process of video frame extraction 710, the bullet screen information 741 in the structured information 740 of the video sample can be analyzed first, and the descriptive words related to the items appearing in the video sample (such as key descriptive words like "clothes", "coat", "good-looking", "beautiful", etc.) can be extracted. Then, the video sample is frame-extracted according to the time interval where these descriptive words are located (i.e., the previous second time interval) to obtain a video sample frame interval. Since these descriptive words in the bullet screen information can reflect the real-time feedback and needs of the audience for the video content, by extracting these descriptive words from the bullet screen information and analyzing the time points when these descriptive words appear in the video, not only can the focus of the audience's interest be found more accurately, but also the interference of redundant pictures in the video sample can be reduced, which is beneficial to improving the accuracy of commodity push. In addition, in the case where the video sample has no bullet screen information or the number of bullet screen information is small, the video sample can be frame-extracted in a uniform frame extraction manner to obtain a video sample frame interval.
[0134] During the processing of clustering and duplicate removal 720, the obtained video sample frame intervals can be subjected to clustering and duplicate removal, thereby deleting some duplicate approximate frames and reducing the processing pressure for subsequent deep learning analysis. In one embodiment, when performing clustering and duplicate removal on the video sample frames in the video sample frame intervals, the basic image features can be used as the benchmark for calculating the relative distance for clustering and duplicate removal, and a specified frame sequence difference can be supplemented as the measure of distance at the same time. Specifically, in each video sample frame interval, first perform frame clustering based on similarity to obtain multiple video sample frame sets, and then in each video sample frame set, retain one candidate video sample frame and delete other candidate video sample frames in the video sample frame set; alternatively, in each video sample frame interval, first perform frame interval segmentation according to the preset frame sequence difference to obtain multiple video sample frame sub-intervals, and then in each video sample frame sub-interval, retain one candidate video sample frame and delete other candidate video sample frames in the video sample frame sub-interval; or, in each video sample frame interval, first perform frame interval segmentation according to the preset frame sequence difference to obtain multiple video sample frame sub-intervals, and then in each video sample frame sub-interval, calculate the similarity between every two adjacent candidate video sample frames. If all similarities are greater than or equal to the similarity threshold, retain one candidate video sample frame and delete other candidate video sample frames in the video sample frame sub-interval. If there is a case where the similarity is less than the similarity threshold among all similarities, retain the two candidate video sample frames with similarities less than the similarity threshold and delete other candidate video sample frames in the video sample frame sub-interval.
[0135] During the processing of the basic filtering 730, the video sample frame intervals that have been clustered and de-duplicated can be first subjected to frame screening based on clarity to delete video sample frames with poor clarity and insufficient brightness, and then the video sample frames in the video sample frame intervals can be further screened using the structured information 740 of the video samples. Specifically, the face recognition result information 742 in the structured information 740 can be first used to screen the video sample frames in the video sample frame intervals to obtain video sample frames with an object face (such as a star face). For example, assuming that the face recognition result information 742 includes star face information, then the video sample frames without star face information can be deleted according to the star face information. Then, the human body detection result information 743 in the structured information 740 can be used to screen the video sample frames with an object face (such as a star face) based on the person's appearance range. Specifically, the position (such as the hip position) and proportion of the human body key points in the human body detection result information 743 can be used to judge the corresponding person's appearance range. For example, it can be judged whether the person in the video sample frame shows the hips. If the person in the video sample frame shows the hips, it means that the person's appearance range meets the requirements. Therefore, the video sample frames with the hips showing can be determined as the target video sample frames.
[0136] In one embodiment, when using the item detection and product retrieval in the second part 620 to obtain the product information of the product samples, the items appearing in the target video sample frames can be first identified, and then the corresponding label information can be generated for the items appearing in the target video sample frames. Then, the product retrieval can be performed using the label information and the images of the items appearing in the target video sample frames. Refer to Figure 8 shown Figure 8 Exemplarily, a flowchart of the product information retrieval is given. In Figure 8 it, the process of the product information retrieval mainly includes three processing processes: item identification 810, label generation 820, and multi-modal product retrieval 830.
[0137] During the processing of the item identification 810, the items appearing in the target video sample frames can be first identified to obtain the types and position information of the items appearing therein (such as Figure 8 "scarf, coordinate 1", "top, coordinate 2", etc. in Figure 8"Scarf, coordinate 1, color 1,...", "Top, coordinate 2, color 1,...", etc.), then, the image of the item that appears can be cropped from the target video sample frame according to the position information of the item that appears, and the multi-modal retrieval model can be called to retrieve product information based on the complete label information and the image of the item that appears, and the product information of multiple candidate product samples can be obtained. At this time, the complete label information and the image of the item that appears can be used to calculate the similarity with the product information of each candidate product sample, so that the product information of the candidate product sample with the largest calculated similarity can be used as the product information of the product sample most similar to the item that appears.
[0138] Step 330: Push the product information of the target product to the first object on the screen interface of the current video playback.
[0139] In one embodiment, after obtaining the product information of the target product, the product information of the target product can be pushed to the first object on the screen interface of the current video playback in the form of a floating window or bullet screen, or, the product information of the target product can be pushed to the first object by displaying a product push interface area on the screen interface of the current video playback, which is not specifically limited here. By pushing the product information of the target product to the first object on the screen interface of the current video playback, not only does it not require switching the current video playback interface to the shopping platform, but also it does not require product retrieval of the target product on the shopping platform. Therefore, it can reduce the cost for the first object to obtain product information that meets its purchase preferences and purchase needs without affecting the viewing integrity of the first object for the video content. In addition, since the target product pushed on the screen interface of the current video playback is retrieved according to the target text information input by the first object, it can better meet the shopping needs of the first object when watching the video, enabling the first object to more easily obtain the product information of the corresponding product according to its purchase preferences or purchase needs, thereby effectively improving the shopping satisfaction of the first object.
[0140] In this embodiment, through the product push method including the previous steps 310 to 330, during the video playback, when the target text information input by the first object is received, the current video playback content is subjected to item recognition according to the target text information to obtain the target item information. Since the target item information is obtained by performing item recognition on the current video playback content according to the target text information input by the first object, the target item information can meet the preferences and needs of the first object. Then, according to the target item information, product information retrieval is performed in the pre-set product information library to obtain the product information of the target product, and in the screen interface of the current video playback, the product information of the target product is pushed to the first object. Since the product information of the target product is obtained by performing product information retrieval in the pre-set product information library according to the target item information, the target product can better meet the purchase preferences and purchase needs of the first object. In addition, by pushing the product information of the target product to the first object in the screen interface of the current video playback, not only does it not need to switch the current video playback interface to the shopping platform, but also it does not need to perform product retrieval on the target product in the shopping platform. Therefore, it can reduce the cost for the first object to obtain product information that meets its purchase preferences and purchase needs without affecting the viewing integrity of the video content by the first object. In addition, since the target product pushed in the screen interface of the current video playback is retrieved according to the target text information input by the first object, it can better meet the shopping needs of the first object when watching the video, enabling the first object to more easily obtain the product information of the corresponding product according to its purchase preferences or purchase needs, thereby effectively improving the shopping satisfaction of the first object.
[0141] In one embodiment, when the screen interface includes a video playback interface area and a product push interface area, and the current video playback content is displayed in the video playback interface area and the product information of the target product is displayed in the product push interface area, when it is detected that the first object triggers the product information of the target product (for example, the first object clicks on the product link of the target product), the current video playback content can be kept displayed in the video playback interface area, and the product purchase page of the target product can be jump-displayed in the product push interface area. When it is detected that the first object completes the product purchase for the target product on the product purchase page, the product push interface area can be not displayed. By this way of real-time interaction during video viewing to obtain product push information, not only can the first object more easily find the products it is interested in during the video viewing process, but also when the first object makes a product purchase, it will not affect its viewing of the current video playback content, and it can also simplify the shopping process of the first object and reduce the operation cost for the first object to obtain product information that meets its purchase preferences and purchase needs, thereby effectively improving the shopping experience of the first object during video viewing.
[0142] In one embodiment, after obtaining the product information of the target product, when it is detected that the first object triggers the product information of the target product (for example, the first object clicks on the product link of the target product), the currently playing video can be paused, and the screen interface can be jumped to display the product purchase page of the target product. When it is detected that the first object completes the purchase of the target product on the product purchase page, the screen interface can be jumped to display the original video playback page, and the current video can be continued to be played. Pausing the currently playing video before jumping the screen interface to display the product purchase page of the target product and continuing to play the current video after jumping the screen interface to display the original video playback page can not affect the viewing integrity of the first object for the current video playback content. Thus, it is possible to improve the shopping satisfaction of the first object while enabling the first object to more easily obtain the product information of the corresponding product according to his / her purchase preferences or purchase needs.
[0143] The following uses a specific example to detail the specific process of the product push method provided by the embodiments of the present application.
[0144] In a specific example, during the process of a user watching a video through a terminal, when the user watches a video such as Figure 9When playing the video content in the video playback interface area 910 of the screen interface, the user posts barrage information 911 (i.e., the target text information) such as "So beautiful", "The dress is so beautiful", or "I like this dress" regarding the dress worn by the female lead in the movie or TV drama being watched. At this time, the terminal will send the barrage information 911 posted by the user to the server. When the server receives the barrage information 911 posted by the user, the server first determines the first time interval where the barrage information 911 is located, and based on the first time interval, determines multiple candidate video frames corresponding to the current video playback content. Then, it performs frame deduplication and clarity screening on these candidate video frames to obtain multiple first video frames that meet the clarity requirements. Then, it performs frame screening for the face of the female lead in the movie or TV drama (i.e., the target face) on these first video frames to obtain multiple second video frames including the face of the female lead in the movie or TV drama. Next, it performs frame screening for the body area range of the female lead in the movie or TV drama on these second video frames to obtain third video frames that meet the body area range requirements (such as requiring the upper body area above the waist or the upper body area above the hips, etc.), and uses the obtained third video frames as the target video frames corresponding to the current video playback content. At this time, the server performs semantic recognition on the barrage information 911 to obtain the target text semantic information of the barrage information 911, and performs item recognition on the target video frames to obtain multiple candidate item information in the target video frames, such as the image information and its label information of the dress worn by the female lead in the movie or TV drama, the image information and its label information of the bag carried by the female lead in the movie or TV drama, the image information and its label information of the lamps in the environment where the female lead in the movie or TV drama is located, etc. Then, among these candidate item information, it determines the target item information that matches the target text semantic information. Since the target text semantic information of the barrage information 911 expresses the preference and praise for the dress worn by the female lead in the movie or TV drama, the server can determine that the target item information that matches the target text semantic information among these candidate item information is the image information and its label information of the dress worn by the female lead in the movie or TV drama. At this time, the server can perform a search for commodity information in the pre-set commodity information library based on the image information and its label information of this dress (i.e., the target item information) to obtain the commodity information of this dress (i.e., the commodity information of the target commodity). Among them, the commodity information of this dress can include the image information of this dress, a title description (such as the same style floral dress, etc.), the price, and the purchase link, etc. Then, the server sends the commodity information of this dress along with the video frames of the current playing video to the terminal, so that through the terminal, in the commodity push interface area 920 of the current video playback screen interface, for example Figure 9 in the commodity push interface area 920 of the screen interface shown, the commodity information of this dress is pushed to the user.
[0145] In another specific example, during the process of a user watching a video through a terminal, when the user watches the video content played in the video playback interface area 1010 of the screen interface as shown in Figure 10 , the user inputs the description information of the commodity to be purchased, such as "suit", through the search bar 1030 in the screen interface and clicks the search button 1031 in the search bar 1030. At this time, the terminal sends the description information of the commodity input by the user to the server. When the server receives the description information of the commodity, the server first determines the first time interval where the description information of the commodity is located, and determines multiple candidate video frames corresponding to the current video playback content according to the first time interval. Then, frame duplication removal and clarity screening are performed on these candidate video frames to obtain multiple first video frames meeting the clarity requirements. Then, frame screening for the human faces (i.e., target human faces) of the main characters such as the female lead in the movie or TV drama is performed on these first video frames to obtain multiple second video frames including the target human faces. Then, frame screening for the human body area range corresponding to the target human face is performed on these second video frames to obtain third video frames meeting the requirements of the human body area range (for example, requiring the upper body area above the waist or the upper body area above the hips, etc.), and the obtained third video frames are used as the target video frames corresponding to the current video playback content. At this time, the server performs semantic recognition on the description information of the commodity to obtain the target text semantic information of the description information of the commodity, and performs item recognition on the target video frames to obtain multiple candidate item information in the target video frames, such as the image information and its label information of the suit worn by the female lead in the movie or TV drama, the image information and its label information of the clothes worn by the male supporting role in the movie or TV drama, the image information and its label information of the clothes worn by passers-by in the movie or TV drama, etc. Then, among these candidate item information, the target item information matching the target text semantic information is determined. Since the target text semantic information of the description information of the commodity is the description of the suit worn by the main character in the movie or TV drama, the server can determine that the target item information matching the target text semantic information among these candidate item information is the image information and its label information of the suit worn by the female lead in the movie or TV drama. At this time, the server can perform a commodity information search in the preset commodity information database according to the image information and its label information of the suit (i.e., the target item information) to obtain the commodity information of the suit (i.e., the commodity information of the target commodity). Among them, the commodity information of the suit can include the image information of the suit, the title description (such as the same style blue suit jacket, etc.), the price, and the purchase link, etc. Then, the server sends the commodity information of the suit along with the video frames of the current played video to the terminal, so as to push the commodity information of the suit to the user in the commodity push interface area 1020 of the screen interface as shown in Figure 10 .
[0146] In addition, as the video continues to play, since the description information of the product that the user wants to purchase (i.e., "suit") remains in the search bar 1030 on the screen interface, then when the video continues to play to the next scene content, for example Figure 11 when playing the video content shown in the video playback interface area 1110 of the screen interface, the server can re-determine multiple candidate video frames corresponding to the current video playback content according to the first time interval where the description information of the product is located, then perform frame deduplication and clarity screening on these candidate video frames to obtain multiple new first video frames that meet the clarity requirements, and then perform frame screening on these new first video frames for the faces of the main characters in the movie or TV drama (i.e., the target faces) to obtain multiple new second video frames including the target faces. Then, perform frame screening on these new second video frames for the range of the human body area corresponding to the target face to obtain new third video frames that meet the requirements of the human body area range (such as requiring the upper body area above the waist or the upper body area above the hips, etc.), and use the obtained new third video frames as the target video frames corresponding to the current video playback content. At this time, the server performs semantic recognition on the description information of the product to obtain the target text semantic information of the description information of the product, and performs object recognition on the target video frames to obtain multiple candidate object information in the target video frames, such as the image information and its label information of the new suit worn by the female lead in the movie or TV drama, the image information and its label information of the clothes worn by the male supporting role in the movie or TV drama, etc. Then, among these candidate object information, determine the target object information that matches the target text semantic information. Since the target text semantic information of the description information of the product is still the description of the suit worn by the main characters in the movie or TV drama, the server can determine that the target object information that matches the target text semantic information among these candidate object information is the image information and its label information of the new suit worn by the female lead in the movie or TV drama. At this time, the server can perform a product information search in the pre-set product information library according to the image information and its label information of this new suit (i.e., the target object information) to obtain the product information of this new suit (i.e., the product information of the target product). Among them, the product information of this new suit can include the image information of this new suit, the title description (such as the same style blue suit jacket, etc.), the price, and the purchase link, etc. Then, the server sends the product information of this new suit along with the video frames of the currently playing video to the terminal, so that through the terminal, in the screen interface of the currently playing video, for example Figure 11 in the product push interface area 1120 of the screen interface shown, push the product information of this new suit to the user.
[0147] Refer to Figure 12 shown Figure 12It is a detailed flowchart of a product push method provided by a specific example. In Figure 12 this method, the product push method may include but is not limited to steps 1201 to 1221.
[0148] Step 1201: Obtain a video sample and obtain multiple barrage information of the video sample.
[0149] Step 1202: Perform semantic recognition on each barrage information to obtain multiple descriptive words related to the items appearing in the video sample.
[0150] Step 1203: Determine the second time interval where each descriptive word is located, and perform frame interval screening on the video sample according to multiple second time intervals to obtain multiple video sample frame intervals.
[0151] Step 1204: Perform frame deduplication and clarity screening on each video sample frame interval to obtain multiple first video sample frames that meet the clarity requirements in each video sample frame interval.
[0152] In one embodiment, the video sample frame interval may include multiple candidate video sample frames. In this case, when performing frame deduplication on each video sample frame interval, frame clustering based on similarity can be performed in each video sample frame interval to obtain multiple video sample frame sets, and then in each video sample frame set, one of the candidate video sample frames is retained, and other candidate video sample frames in the video sample frame set are deleted. Alternatively, when performing frame deduplication on each video sample frame interval, the video sample frame interval can also be segmented according to a preset frame sequence difference in each video sample frame interval to obtain multiple video sample frame sub-intervals; then in each video sample frame sub-interval, one of the candidate video sample frames is retained, and other candidate video sample frames in the video sample frame sub-interval are deleted. In addition, when performing frame deduplication on each video sample frame interval, the video sample frame interval can also be segmented according to a preset frame sequence difference in each video sample frame interval to obtain multiple video sample frame sub-intervals; then in each video sample frame sub-interval, the similarity between every two adjacent candidate video sample frames is calculated; when all similarities are greater than or equal to the similarity threshold, one of the candidate video sample frames is retained, and other candidate video sample frames in the video sample frame sub-interval are deleted; when there is a case where all similarities are less than the similarity threshold, two candidate video sample frames with similarities less than the similarity threshold are retained, and other candidate video sample frames in the video sample frame sub-interval are deleted.
[0153] Step 1205: Perform frame screening on multiple first video sample frames according to a preset object face to obtain multiple second video sample frames including the object face.
[0154] In one embodiment, the pre-set object face can be a star face or the face of the main character in the video sample (such as the faces of the male and female protagonists or the male and female supporting characters in the video sample, etc.), which is not specifically limited here.
[0155] Step 1206: Perform frame screening on multiple second video sample frames for the human body area range corresponding to the object face to obtain multiple target video sample frames that meet the requirements of the human body area range.
[0156] In one embodiment, the human body area range can be the upper body area including above the waist, the upper body area including above the hips, or the entire human body area including both the upper body and the lower body, which can be appropriately selected according to the actual application situation and is not specifically limited here.
[0157] Step 1207: Perform item recognition on each target video sample frame to obtain the target item sample information in each target video sample frame.
[0158] Step 1208: Perform a commodity information retrieval on each target item sample information to obtain the commodity information of the commodity sample corresponding to each target item sample information.
[0159] In one embodiment, the commodity information can include commodity images, commodity labels, commodity links, etc.
[0160] Step 1209: Construct a commodity information library based on the commodity information of the video sample and the commodity sample.
[0161] In one embodiment, when constructing a commodity information library based on the commodity information of the video sample and the commodity sample, the video ID of the video sample can be obtained first, and then the video ID of the video sample and the commodity information of the commodity sample are used to construct the commodity information library. In this way, the video sample and the commodity information library can be associated through the video ID, so that each video sample can correspond to a commodity information library. In addition, the commodity information library can also include star face information or the face information of the main character in the video sample, so the efficiency of retrieving the same-style commodities of the video character through the commodity information library can be improved.
[0162] Step 1210: During the video playback, detect whether the target text information input by the first object is received. If so, execute step 1211; if not, execute step 1210.
[0163] In one embodiment, the first object can be the user watching the video. The target text information input by the first object can be the bullet screen information related to the items appearing in the video playback content posted by the user, or it can be the item description information related to the items appearing in the video playback content input by the user in the search bar, which is not specifically limited here.
[0164] Step 1211: Determine the first time interval where the target text information is located, and determine multiple candidate video frames corresponding to the current video playback content according to the first time interval.
[0165] Step 1212: Perform frame deduplication and clarity screening on the multiple candidate video frames to obtain multiple first video frames that meet the clarity requirements.
[0166] Step 1213: Perform frame screening on the multiple first video frames for the target face to obtain multiple second video frames including the target face.
[0167] In one embodiment, the target face can be the star face information included in the commodity information database or the face information of the main character in the video sample.
[0168] Step 1214: Perform frame screening on the multiple second video frames for the human body area range corresponding to the target face to obtain third video frames that meet the human body area range requirements, and use the third video frames as the target video frames corresponding to the current video playback content.
[0169] In one embodiment, the human body area range corresponding to the target face can be the upper body area corresponding to the target face including above the waist, the upper body area corresponding to the target face including above the hips, or the entire human body area corresponding to the target face including both the upper body and the lower body, and can be appropriately selected according to the actual application situation, and no specific limitation is made here.
[0170] In one embodiment, when performing frame deduplication on the multiple candidate video frames, frame clustering based on similarity can be performed on the multiple candidate video frames to obtain multiple video frame sets, and then in each video frame set, one of the candidate video frames is retained, and other candidate video frames in the video frame set are deleted. Or, when performing frame deduplication on the multiple candidate video frames, the multiple candidate video frames can also be frame-separated according to a preset frame sequence difference to obtain multiple video frame intervals; then in each video frame interval, one of the candidate video frames is retained, and other candidate video frames in the video frame interval are deleted. In addition, when performing frame deduplication on the multiple candidate video frames, the multiple candidate video frames can also be frame-separated according to a preset frame sequence difference to obtain multiple video frame intervals; then in each video frame interval, the similarity between every two adjacent candidate video frames is calculated; when all similarities are greater than or equal to the similarity threshold, one of the candidate video frames is retained, and other candidate video frames in the video frame interval are deleted; when there is a case where all similarities are less than the similarity threshold, two candidate video frames with similarities less than the similarity threshold are retained, and other candidate video frames in the video frame interval are deleted.
[0171] Step 1215: Perform item recognition on the target video frame according to the target text information to obtain the target item information.
[0172] In one embodiment, when performing item recognition on the target video frame according to the target text information to obtain the target item information, the semantic recognition of the target text information can be performed first to obtain the target text semantic information, then item recognition is performed on the target video frame to obtain multiple candidate item information in the target video frame, and then among the multiple candidate item information, the target item information that matches the target text semantic information is determined.
[0173] Step 1216: Retrieve the product information of the target product in the pre-set product information library according to the target item information to obtain the product information of the target product.
[0174] Step 1217: Push the product information of the target product to the first object on the screen interface of the current video playback.
[0175] Step 1218: Detect whether the first object has triggered the product information of the target product. If so, execute step 1219; if not, execute step 1218.
[0176] Step 1219: Display the product purchase page of the target product.
[0177] Step 1220: Detect whether the first object has completed the purchase of the target product on the product purchase page. If so, execute step 1221; if not, execute step 1220.
[0178] Step 1221: Do not display the product purchase page of the target product and execute step 1210.
[0179] In one embodiment, after pushing the product information of the target product to the first object, in the case where the screen interface includes a video playback interface area and a product push interface area, the video playback interface area displays the current video playback content, and the product push interface area displays the product information of the target product. If it is detected that the first object triggers the product information of the target product, the current video playback content can be kept displayed in the video playback interface area, and the product purchase page of the target product can be jump-displayed in the product push interface area. When it is detected that the first object has completed the purchase of the target product on the product purchase page, the product push interface area is not displayed.
[0180] In another embodiment, after pushing the product information of the target product to the first object, if it is detected that the first object triggers the product information of the target product, the currently playing video can be paused, and the screen interface can be jumped to display the product purchase page of the target product. When it is detected that the first object has completed the purchase of the target product on the product purchase page, the screen interface is jumped to display the original video playing page, and the current video is continued to be played.
[0181] In this embodiment, through the product pushing method of the above steps 1201 to 1221, during the video playing process, when the target text information input by the user is received, the item recognition is performed on the current video playing content according to the target text information to obtain the target item information. Since the target item information is obtained by performing item recognition on the current video playing content according to the target text information input by the user, the target item information can meet the user's preferences and needs. Then, according to the target item information, the product information is retrieved in the preset product information library to obtain the product information of the target product, and the product information of the target product is pushed to the user on the screen interface of the current video playing. Since the product information of the target product is obtained by retrieving the product information in the preset product information library according to the target item information, the target product can better meet the user's purchase preferences and purchase needs. In addition, by pushing the product information of the target product to the user on the screen interface of the current video playing, not only does it not need to switch the current video playing interface to the shopping platform, but also it does not need to retrieve the product on the shopping platform. Therefore, the cost for the user to obtain the product information that meets their purchase preferences and purchase needs can be reduced without affecting the viewing integrity of the video content for the user. In addition, since the target product pushed on the screen interface of the current video playing is retrieved according to the target text information input by the user, the shopping needs of the user when watching the video can be better met, enabling the user to more easily obtain the product information of the corresponding product according to their purchase preferences or purchase needs, thereby effectively improving the shopping satisfaction of the user.
[0182] The following uses some actual examples to illustrate the application scenarios of the embodiments of the present application.
[0183] It should be noted that the product pushing method provided by the embodiments of the present application can be applied to different application scenarios such as obtaining the same model products of the main characters in the video through bullet screens and obtaining the same model products of the main characters in the video through search terms. The following takes the scenarios of obtaining the same model products of the main characters in the video through bullet screens and obtaining the same model products of the main characters in the video through search terms as examples for illustration.
[0184] Scenario 1
[0185] The product push method provided by the embodiments of this application can be applied to scenarios where the same products of the main characters in a video are obtained through bullet comments. For example, during the process of a user watching a video through a terminal, when the user posts bullet comment information such as "so beautiful", "the clothes are really beautiful", or "like this piece of clothing" for the clothes worn by the female lead in the watched TV drama, the terminal will send the bullet comment information posted by the user to the server. After receiving the bullet comment information, the server performs item recognition on the current video playback content according to the bullet comment information to obtain target item information, and then retrieves product information in a preset product information library according to the target item information to obtain the product information of the target product. Then, the server pushes the product information of the target product to the user on the screen interface of the current video playback. Since the product information of the target product pushed to the user is obtained by retrieving products according to the bullet comment information posted by the user, the target product can meet the user's preferences and needs, can better meet the user's shopping needs when watching the video, enables the user to more easily obtain the product information of the corresponding product according to their purchase preferences or purchase needs, and thus can effectively improve the user's shopping satisfaction.
[0186] Scenario 2
[0187] The product push method provided by the embodiments of this application can also be applied to scenarios where the same products of the main characters in a video are obtained through search terms. For example, during the process of a user watching a video through a terminal, when the user enters a search term for the product they want to buy, such as the description information of the product, through the search bar on the screen interface and clicks the search button in the search bar, the terminal will send the search term entered by the user to the server. After receiving the search term, the server performs item recognition on the current video playback content according to the search term to obtain target item information, and then retrieves product information in a preset product information library according to the target item information to obtain the product information of the target product. Then, the server pushes the product information of the target product to the user on the screen interface of the current video playback. Since the product information of the target product pushed to the user is obtained by retrieving products according to the search term entered by the user, the target product can meet the user's preferences and needs, can better meet the user's shopping needs when watching the video, enables the user to more easily obtain the product information of the corresponding product according to their purchase preferences or purchase needs, and thus can effectively improve the user's shopping satisfaction.
[0188] It can be understood that although the steps in the above various flowcharts are sequentially shown according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0189] Referring to Figure 13 , the embodiment of the present application also discloses a commodity push device. The commodity push device 1300 can implement the commodity push method in the previous embodiment. The commodity push device 1300 includes:
[0190] An item recognition unit 1310, configured to receive target text information input by a first object during video playback, and perform item recognition on the current video playback content according to the target text information to obtain target item information;
[0191] A commodity retrieval unit 1320, configured to retrieve commodity information in a preset commodity information database according to the target item information to obtain the commodity information of the target commodity;
[0192] A commodity push unit 1330, configured to push the commodity information of the target commodity to the first object on the screen interface of the current video playback.
[0193] In one embodiment, the item recognition unit 1310 is further configured to:
[0194] Determine a target video frame corresponding to the current video playback content according to the target text information;
[0195] Perform item recognition on the target video frame according to the target text information to obtain target item information.
[0196] In one embodiment, the item recognition unit 1310 is further configured to:
[0197] Determine a first time interval where the target text information is located, and determine multiple candidate video frames corresponding to the current video playback content according to the first time interval;
[0198] Perform frame screening processing on the multiple candidate video frames to obtain a target video frame corresponding to the current video playback content.
[0199] In one embodiment, the item recognition unit 1310 is further configured to:
[0200] Perform frame deduplication and clarity screening on multiple candidate video frames to obtain multiple first video frames that meet the clarity requirements;
[0201] Perform frame screening for the target face on multiple first video frames to obtain multiple second video frames including the target face;
[0202] Perform frame screening for the human body region range corresponding to the target face on multiple second video frames to obtain third video frames that meet the human body region range requirements;
[0203] Use the third video frame as the target video frame corresponding to the current video playback content.
[0204] In one embodiment, the item recognition unit 1310 is further configured to:
[0205] Perform frame clustering based on similarity on multiple candidate video frames to obtain multiple video frame sets. In each video frame set, retain one of the candidate video frames and delete the other candidate video frames in the video frame set;
[0206] Alternatively, perform frame separation on multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; in each video frame interval, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval;
[0207] Alternatively, perform frame separation on multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; in each video frame interval, calculate the similarity between every two adjacent candidate video frames; when all similarities are greater than or equal to the similarity threshold, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval; when there is a case where any of the similarities is less than the similarity threshold, retain the two candidate video frames with similarities less than the similarity threshold and delete the other candidate video frames in the video frame interval.
[0208] In one embodiment, the screen interface includes a video playback interface area and a commodity push interface area. The video playback interface area displays the current video playback content, and the commodity push interface area displays the commodity information of the target commodity; the commodity push device 1300 further includes:
[0209] A first page jump unit, configured to, when detecting that a first object triggers the commodity information of the target commodity, keep displaying the current video playback content in the video playback interface area and jump to display the commodity purchase page of the target commodity in the commodity push interface area;
[0210] A first page processing unit, configured to, when detecting that the first object has completed the commodity purchase for the target commodity on the commodity purchase page, not display the commodity push interface area.
[0211] In one embodiment, the product push device 1300 further includes:
[0212] A second page jump unit, configured to pause the currently playing video when detecting that a first object triggers the product information of a target product, and jump and display the screen interface as the product purchase page of the target product;
[0213] A second page processing unit, configured to jump and display the screen interface as the original video playing page and continue playing the current video when detecting that the first object has completed the purchase of the target product on the product purchase page.
[0214] In one embodiment, the product push device 1300 further includes a product information library construction unit, and the product information library construction unit is used for:
[0215] Obtain video samples, perform frame interval screening on the video samples to obtain a plurality of video sample frame intervals;
[0216] Perform frame screening processing on the plurality of video sample frame intervals respectively to obtain a plurality of target video sample frames;
[0217] Perform object recognition on each target video sample frame to obtain target object sample information in each target video sample frame;
[0218] Perform product information retrieval on each target object sample information to obtain the product information of the product samples corresponding to each target object sample information;
[0219] Construct a product information library according to the video samples and the product information of the product samples.
[0220] In one embodiment, the product information library construction unit is further used for:
[0221] Obtain a plurality of barrage information of the video samples, perform semantic recognition on each barrage information to obtain a plurality of descriptive words related to the objects appearing in the video samples;
[0222] Determine the second time interval where each descriptive word is located, and perform frame interval screening on the video samples according to the plurality of second time intervals to obtain a plurality of video sample frame intervals.
[0223] In one embodiment, the product information library construction unit is further used for:
[0224] Perform frame deduplication and clarity screening on each video sample frame interval to obtain a plurality of first video sample frames meeting the clarity requirements in each video sample frame interval;
[0225] Perform frame screening on the plurality of first video sample frames according to the preset object face to obtain a plurality of second video sample frames including the object face;
[0226] Perform frame screening on multiple second video sample frames for the human body region range corresponding to the object face to obtain multiple target video sample frames that meet the requirements of the human body region range.
[0227] In one embodiment, the video sample frame interval includes multiple candidate video sample frames; the commodity information library construction unit is further configured to:
[0228] In each video sample frame interval, perform frame clustering based on similarity to obtain multiple video sample frame sets. In each video sample frame set, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame set;
[0229] Alternatively, in each video sample frame interval, perform frame interval segmentation based on a preset frame sequence difference to obtain multiple video sample frame sub-intervals; in each video sample frame sub-interval, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame sub-interval;
[0230] Alternatively, in each video sample frame interval, perform frame interval segmentation based on a preset frame sequence difference to obtain multiple video sample frame sub-intervals; in each video sample frame sub-interval, calculate the similarity between every two adjacent candidate video sample frames; when all similarities are greater than or equal to the similarity threshold, retain one of the candidate video sample frames and delete the other candidate video sample frames in the video sample frame sub-interval; when there is a case where any of the similarities is less than the similarity threshold, retain the two candidate video sample frames with similarities less than the similarity threshold and delete the other candidate video sample frames in the video sample frame sub-interval.
[0231] It should be noted that since the commodity push device 1300 in this embodiment can implement the commodity push method in the previous embodiment, the commodity push device 1300 in this embodiment and the commodity push method in the previous embodiment have the same technical principle and the same beneficial effects. To avoid repetition of content, it will not be elaborated here.
[0232] Refer to Figure 14 , this application embodiment also discloses an electronic device, and this electronic device 1400 includes:
[0233] At least one processor 1401;
[0234] At least one memory 1402 for storing at least one program;
[0235] When at least one program is executed by at least one processor 1401, the commodity push method as described above is implemented.
[0236] The embodiments of the present application also disclose a computer-readable storage medium, in which a computer program executable by a processor is stored. When the computer program executable by the processor is executed by the processor, it is used to implement the commodity push method as described above.
[0237] The embodiments of the present application also disclose a computer program product, including a computer program or computer instructions. The computer program or computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the electronic device executes the commodity push method as described above.
[0238] In the description of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0239] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0240] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.
[0241] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0242] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0243] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0244] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0245] Regarding the step numbers in the above method embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
Claims
1. A commodity push method, characterized in that Including the following steps: During video playback, receive the target text information input by the first object, and perform object recognition on the current video playback content according to the target text information to obtain target object information; Retrieve product information in a preset product information database according to the target object information to obtain the product information of the target product; Push the product information of the target product to the first object on the screen interface of the current video playback.
2. The method according to claim 1, wherein The performing object recognition on the current video playback content according to the target text information to obtain target object information includes: Determine the target video frame corresponding to the current video playback content according to the target text information; Perform object recognition on the target video frame according to the target text information to obtain target object information.
3. The method according to claim 2, characterized in that The determining the target video frame corresponding to the current video playback content according to the target text information includes: Determine the first time interval where the target text information is located, and determine multiple candidate video frames corresponding to the current video playback content according to the first time interval; Perform frame screening processing on the multiple candidate video frames to obtain the target video frame corresponding to the current video playback content.
4. The method according to claim 3, characterized in that The performing frame screening processing on the multiple candidate video frames to obtain the target video frame corresponding to the current video playback content includes: Perform frame deduplication and clarity screening on the multiple candidate video frames to obtain multiple first video frames meeting the clarity requirements; Perform frame screening on the multiple first video frames for the target human face to obtain multiple second video frames including the target human face; Perform frame screening on the multiple second video frames for the human body area range corresponding to the target human face to obtain third video frames meeting the human body area range requirements; Use the third video frame as the target video frame corresponding to the current video playback content.
5. The method according to claim 4, wherein The process of performing frame deduplication on the multiple candidate video frames includes: Perform frame clustering based on similarity on the multiple candidate video frames to obtain multiple video frame sets. In each video frame set, retain one of the candidate video frames and delete the other candidate video frames in the video frame set; Alternatively, perform frame separation on the multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; in each video frame interval, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval; Alternatively, perform frame separation on the multiple candidate video frames according to a preset frame sequence difference to obtain multiple video frame intervals; in each video frame interval, calculate the similarity between every two adjacent candidate video frames; when all the similarities are greater than or equal to the similarity threshold, retain one of the candidate video frames and delete the other candidate video frames in the video frame interval; when there is a situation where any of the similarities is less than the similarity threshold, retain the two candidate video frames with similarities less than the similarity threshold and delete the other candidate video frames in the video frame interval.
6. The method according to claim 1, characterized in that The screen interface includes a video playback interface area and a product push interface area. The video playback interface area displays the current video playback content, and the product push interface area displays the product information of the target product; The method further includes: When it is detected that the first object triggers the product information of the target product, the current video playback content is kept displayed in the video playback interface area, and the product purchase page of the target product is jump-displayed in the product push interface area; When it is detected that the first object has completed the product purchase of the target product on the product purchase page, the product push interface area is not displayed.
7. The method according to claim 1, characterized in that The method further includes: When it is detected that the first object triggers the product information of the target product, the currently playing video is paused, and the screen interface is jump-displayed as the product purchase page of the target product; When it is detected that the first object has completed the product purchase of the target product on the product purchase page, the screen interface is jump-displayed as the original video playback page, and the current video continues to play.
8. The method according to claim 1, wherein The product information library is constructed according to the following steps: Obtain video samples, perform frame interval screening on the video samples to obtain multiple video sample frame intervals; Perform frame screening processing on each of the multiple video sample frame intervals to obtain multiple target video sample frames; Perform object recognition on each of the target video sample frames to obtain the target object sample information in each of the target video sample frames; Perform product information retrieval on each of the target object sample information to obtain the product information of the product samples corresponding to each of the target object sample information; Construct the product information library according to the video samples and the product information of the product samples.
9. The method according to claim 8, wherein The performing frame interval screening on the video samples to obtain multiple video sample frame intervals includes: Obtain multiple bullet screen information of the video samples, perform semantic recognition on each of the bullet screen information to obtain multiple descriptive words related to the objects appearing in the video samples; Determine the second time interval where each of the descriptive words is located, and perform frame interval screening on the video samples according to the multiple second time intervals to obtain multiple video sample frame intervals.
10. The method according to claim 8, wherein The performing frame screening processing on each of the multiple video sample frame intervals to obtain multiple target video sample frames includes: Perform frame deduplication and clarity screening on each of the video sample frame intervals to obtain multiple first video sample frames that meet the clarity requirements in each of the video sample frame intervals; Perform frame screening on the multiple first video sample frames according to the preset object face to obtain multiple second video sample frames including the object face; Perform frame screening on the multiple second video sample frames for the human body area range corresponding to the object face to obtain multiple target video sample frames that meet the human body area range requirements.
11. The method according to claim 10, wherein The video sample frame interval includes multiple candidate video sample frames; the process of performing frame deduplication on each of the video sample frame intervals includes: In each of the video sample frame intervals, frame clustering based on similarity is performed to obtain multiple video sample frame sets. In each of the video sample frame sets, one of the candidate video sample frames is retained, and the other candidate video sample frames in the video sample frame set are deleted; Alternatively, in each of the video sample frame intervals, frame interval segmentation is performed according to a preset frame sequence difference to obtain multiple video sample frame sub-intervals; in each of the video sample frame sub-intervals, one of the candidate video sample frames is retained, and the other candidate video sample frames in the video sample frame sub-interval are deleted; Alternatively, in each of the video sample frame intervals, frame interval segmentation is performed according to a preset frame sequence difference to obtain multiple video sample frame sub-intervals; in each of the video sample frame sub-intervals, the similarity between every two adjacent candidate video sample frames is calculated; when all the similarities are greater than or equal to a similarity threshold, one of the candidate video sample frames is retained, and the other candidate video sample frames in the video sample frame sub-interval are deleted; when there is a case where at least one of the similarities is less than the similarity threshold, the two candidate video sample frames with similarities less than the similarity threshold are retained, and the other candidate video sample frames in the video sample frame sub-interval are deleted.
12. A commodity pushing device, characterized in that, Comprising: An item recognition unit, configured to receive target text information input by a first object during video playback, and perform item recognition on the current video playback content according to the target text information to obtain target item information; A commodity retrieval unit, configured to retrieve commodity information in a preset commodity information library according to the target item information to obtain the commodity information of a target commodity; A commodity push unit, configured to push the commodity information of the target commodity to the first object on the screen interface of the current video playback.
13. An electronic device, characterized in that, Comprising: At least one processor; At least one memory, configured to store at least one program; When at least one of the programs is executed by at least one of the processors, the commodity push method according to any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium, characterized in that, Wherein there is a computer program executable by a processor, and when the computer program executable by the processor is executed by the processor, it is used to implement the commodity push method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program or computer instructions, characterized in that, The computer program or the computer instruction is stored in a computer-readable storage medium, and a processor of an electronic device reads the computer program or the computer instruction from the computer-readable storage medium, and the processor executes the computer program or the computer instruction, so that the electronic device executes the commodity push method according to any one of claims 1 to 11.