Information pushing method and device and electronic equipment
By acquiring natural language and camera surveillance videos, and utilizing feature data matching and machine learning models, relevant video information is automatically pushed, solving the problem of users having difficulty accurately and promptly judging the occurrence of events in surveillance videos, and improving the accuracy and timeliness of event judgments.
Patent Information
- Application Number
- CN202410494017.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-24
AI Technical Summary
It is difficult for users to determine in a timely and accurate manner whether an event of interest has occurred in a surveillance video, especially when the amount of data is huge and it is difficult to watch the video continuously.
By acquiring natural language and videos generated by camera surveillance, video feature data is extracted, and the event extraction model is trained using machine learning algorithms to automatically match and push relevant video information.
It achieves more accurate and timely push of event-triggered video information, improving the accuracy and timeliness of users' judgment of event occurrence.
Smart Images

Figure CN120835127A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of security and protection, and in particular to an information pushing method and device and electronic equipment. BACKGROUND
[0002] In the related art, in the field of security and protection, a user can retrieve corresponding monitoring video according to collection time, or continuously watch the monitoring video to determine whether an event of interest is included in the monitoring video.
[0003] However, in some cases, the user cannot determine the occurrence time of the event of interest, and thus cannot determine which time interval of the monitoring video to retrieve for watching. In some other cases, the data volume of the monitoring video is huge, and the user cannot continuously and uninterruptedly watch the video. Thus, the user cannot determine whether the event of interest occurs in a timely and accurate manner.
[0004] It can be seen that how to determine whether an event of interest occurs in a more timely and / or accurate manner is a technical problem worthy of attention. SUMMARY
[0005] In view of this, to solve the above-mentioned technical problems, the embodiments of the present application provide an information pushing method, device and electronic equipment.
[0006] In a first aspect, the embodiments of the present application provide an information pushing method, which comprises:
[0007] obtaining natural language and video generated by camera monitoring, wherein the natural language is used to determine a video to be pushed;
[0008] determining feature data of the natural language to obtain first feature data;
[0009] determining a first video based on the video generated by the camera monitoring, and determining whether the first video matches the natural language based on at least two feature data of the first video and the first feature data;
[0010] in a case where the first video matches the natural language, taking the first video as a video to be pushed, and pushing video information of the video to be pushed, wherein the video information represents information of the video to be pushed.
[0011] In one possible implementation, the determination of whether the first video matches the first feature data based on the at least two feature data of the first video comprises:
[0012] determining at least two feature data of image feature data, text feature data and audio feature data of the first video to obtain second feature data;
[0013] determine whether the first video matches the natural language based on the first feature data and the second feature data.
[0014] In one possible implementation, the first video is generated in the following manner:
[0015] extract event frames from videos generated by camera monitoring;
[0016] determine, as the first video, multiple event frames representing the same event, wherein a similarity between the multiple event frames representing the same event is greater than or equal to a preset first threshold, and a similarity between multiple event frames representing different events is less than the preset first threshold.
[0017] In one possible implementation, the extracting event frames from videos generated by camera monitoring comprises:
[0018] extract event frames from videos generated by camera monitoring using an event extraction model; and
[0019] The event extraction model is trained in the following manner:
[0020] obtain a training sample set, wherein a training sample in the training sample set comprises a video, an event time, and an event label;
[0021] train the event extraction model using a machine learning algorithm, wherein the video included in the training sample in the training sample set is used as input data, and the event time and the event label are used as expected output data.
[0022] In one possible implementation, the method further comprises:
[0023] determine a first speed for playing a non-target video segment in the obtained video;
[0024] determine a second speed for playing a target video segment in the obtained video;
[0025] wherein the target video segment is composed of event frames, and the second speed is less than the first speed.
[0026] In one possible implementation, the method further comprises:
[0027] generate a description text of the first video;
[0028] wherein the description text is used for terminal display, and / or the description text is used to determine whether the first video is a search result of a video search request sent by the terminal.
[0029] In a possible implementation, the method further includes:
[0030] determining at least one of text and music matching the first video, to obtain matching information of the first video;
[0031] fusing the matching information with the first video, to obtain a second video;
[0032] in a case where a target operation for the second video is detected, performing the target operation on the second video, wherein the target operation includes at least one of sharing, downloading, storing, and sending.
[0033] In a possible implementation, the method further includes:
[0034] determining a target video frame from the first video, wherein a similarity between the target video frame and a preceding video frame is less than or equal to a preset second threshold, and a similarity between the target video frame and a following video frame is less than or equal to the preset second threshold, the preceding video frame being a previous video frame of the target video frame in the first video, and the following video frame being a next video frame of the target video frame in the first video;
[0035] determining the target video frame as a highlight video frame in the first video.
[0036] In a possible implementation, the method is applied to a first device end; and
[0037] the pushing of the video information of the to-be-pushed video includes:
[0038] obtaining location information of a second device end;
[0039] determining the video information of the to-be-pushed video based on the location information;
[0040] pushing the video information to the second device end.
[0041] In a possible implementation, the determining of the video information of the to-be-pushed video based on the location information includes:
[0042] determining a target location where the camera is located, to obtain a target location;
[0043] determining a distance between a location represented by the location information and the target location, to obtain a target distance;
[0044] determining whether the target distance is greater than or equal to a preset distance threshold;
[0045] In a case that the target distance is greater than or equal to the preset distance threshold, the video information of the to-be-pushed video is determined as first information; wherein the first information indicates a request to control the camera to monitor a target region, and the target region is a region in which the to-be-pushed video is generated.
[0046] In a case that the target distance is less than the preset distance threshold, the video information of the to-be-pushed video is determined as second information; wherein the second information indicates a position of a target region, and the target region is a region in which the to-be-pushed video is generated.
[0047] In a second aspect, an embodiment of the present application provides an information pushing device, and the device comprises:
[0048] An acquisition unit is configured to acquire a natural language and a video generated by camera monitoring, wherein the natural language is used to determine a to-be-pushed video.
[0049] A first determination unit is configured to determine feature data of the natural language to obtain first feature data.
[0050] A second determination unit is configured to determine a first video based on the video generated by the camera monitoring, and determine whether the first video matches the natural language based on at least two kinds of feature data of the first video and the first feature data.
[0051] A pushing unit is configured to, in a case that the first video matches the natural language, push the first video as a to-be-pushed video, and push video information of the to-be-pushed video, wherein the video information indicates information of the to-be-pushed video.
[0052] In a possible implementation, the determination of whether the first video matches the natural language comprises:
[0053] The feature data of the natural language is determined to obtain the first feature data.
[0054] At least two kinds of feature data of image feature data, text feature data and audio feature data of the first video are determined to obtain second feature data.
[0055] Whether the first video matches the natural language is determined based on the first feature data and the second feature data.
[0056] In a possible implementation, the first video is generated in the following manner:
[0057] An event frame is extracted from the video generated by the camera monitoring.
[0058] The extracted multiple event frames representing the same event are determined as the first video, wherein a similarity between the multiple event frames representing the same event is greater than or equal to a preset first threshold, and a similarity between multiple event frames representing different events is less than the preset first threshold.
[0059] In a possible implementation, the extracting the event frames from the video generated by the camera monitoring includes:
[0060] extracting the event frames from the video generated by the camera monitoring by using an event extraction model; and
[0061] The event extraction model is obtained by training in the following manner:
[0062] obtaining a training sample set, wherein a training sample in the training sample set includes a video, an event time, and an event label;
[0063] training the event extraction model by using a machine learning algorithm, taking the video included in the training sample in the training sample set as input data, and taking the event time and the event label as expected output data.
[0064] In a possible implementation, the apparatus further includes:
[0065] a third determination unit, configured to determine a playing speed of a non-target video segment in the obtained video as a first speed;
[0066] a fourth determination unit, configured to determine a playing speed of a target video segment in the obtained video as a second speed;
[0067] wherein the target video segment is composed of event frames, and the second speed is less than the first speed.
[0068] In a possible implementation, the apparatus further includes:
[0069] a generating unit, configured to generate a description text of the first video;
[0070] wherein the description text is used for terminal display, and / or the description text is used for determining whether the first video is a search result of a video search request sent by the terminal.
[0071] In a possible implementation, the apparatus further includes:
[0072] a fifth determination unit, configured to determine at least one of a text and music matched with the first video, to obtain matching information of the first video;
[0073] a fusion unit, configured to fuse the matching information with the first video to obtain a second video;
[0074] a second processing unit, configured to perform a target operation on the second video in a case where the target operation is detected for the second video, wherein the target operation comprises at least one of sharing, downloading, storing, and sending.
[0075] In one possible implementation, the apparatus further includes:
[0076] a sixth determining unit, configured to determine a target video frame from the first video, wherein a similarity between the target video frame and a preceding video frame is less than or equal to a preset second threshold, and a similarity between the target video frame and a following video frame is less than or equal to the preset second threshold, the preceding video frame being a previous video frame of the target video frame in the first video, and the following video frame being a next video frame of the target video frame in the first video;
[0077] a seventh determining unit, configured to determine the target video frame as a highlight video frame in the first video.
[0078] In one possible implementation, the method is applied to a first device end; and
[0079] The pushing of the video information of the to-be-pushed video includes:
[0080] obtaining location information of a second device end;
[0081] determining the video information of the to-be-pushed video based on the location information;
[0082] pushing the video information to the second device end.
[0083] In one possible implementation, the determining of the video information of the to-be-pushed video based on the location information includes:
[0084] determining a target location where the camera is located;
[0085] determining a target distance between a location represented by the location information and the target location;
[0086] determining whether the target distance is greater than or equal to a preset distance threshold;
[0087] in a case where the target distance is greater than or equal to the preset distance threshold, determining the video information of the to-be-pushed video as first information; wherein the first information indicates a request for controlling the camera to monitor a target area, and the target area is an area from which the to-be-pushed video is generated.
[0088] In a case that the target distance is less than the preset distance threshold, it is determined that the video information of the to-be-pushed video represents second information; wherein the second information represents a position of a target area, and the target area is an area in which the to-be-pushed video is generated by monitoring.
[0089] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0090] a memory configured to store a computer program;
[0091] a processor configured to execute the computer program stored in the memory, and when the computer program is executed, the method in any one of the embodiments of the information pushing method of the first aspect of the present application is implemented.
[0092] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method in any one of the embodiments of the information pushing method of the first aspect is implemented.
[0093] In a fifth aspect, an embodiment of the present application provides a computer program product, which comprises computer readable code, and when the computer readable code is executed on a device, the processor in the device implements the method in any one of the embodiments of the information pushing method of the first aspect.
[0094] The information pushing method provided by the embodiments of the present application can obtain natural language and a video generated by camera monitoring, wherein the natural language is used to determine a to-be-pushed video, then feature data of the natural language is determined to obtain first feature data, then a first video is determined based on the video generated by camera monitoring, and whether the first video matches the natural language is determined based on at least two kinds of feature data of the first video and the first feature data, finally, in a case that the first video matches the natural language, the first video is taken as a to-be-pushed video, and video information of the to-be-pushed video is pushed, wherein the video information represents information of the to-be-pushed video. Therefore, the automatic reminding can be realized when an event represented by the natural language is triggered in the video generated by camera monitoring by setting the natural language in advance, the video information triggered by the corresponding event can be pushed more accurately and / or timely, and therefore the accuracy and / or timeliness of judging whether an event concerned by a user or the like occurs is improved. BRIEF DESCRIPTION OF DRAWINGS
[0095] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.
[0096] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those skilled in the field, other drawings can also be obtained based on these drawings without any creative effort.
[0097] One or more embodiments are illustrated by the pictures in the drawings corresponding thereto, which do not constitute a limitation on the embodiments, and elements with the same reference numerals in the drawings represent similar elements, unless otherwise specified. The drawings in the drawings do not constitute a proportional limit.
[0098] Figure 1 A flowchart of an information pushing method provided by an embodiment of the present application;
[0099] Figure 2 A flowchart of another information pushing method provided by an embodiment of the present application;
[0100] Figure 3 An application scenario diagram of an information pushing method provided by an embodiment of the present application;
[0101] Figure 4 A structural diagram of an information pushing device provided by an embodiment of the present application;
[0102] Figure 5 A structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0103] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all the embodiments. It should be noted that: unless otherwise specified, the relative arrangement, numerical expression and values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0104] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent the logical order between them.
[0105] It should also be understood that in the present embodiment, "a plurality of" can mean two or more, and "at least one" can mean one, two or more.
[0106] It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, without explicit limitation or in the context of the opposite indication, it can be understood as one or more in general.
[0107] In addition, the term "and / or" in the present application is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the front and rear associated objects.
[0108] It should also be understood that the description of the various embodiments of the present application focuses on the differences between the various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0109] The following description of at least one example embodiment is merely illustrative in nature and is in no way limiting on the application or its use.
[0110] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered part of the specification when appropriate.
[0111] It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0112] It should be noted that the embodiments and features in the embodiments in the present application can be combined with each other without conflict. In order to facilitate the understanding of the embodiments of the present application, the present application will be described in detail below with reference to the drawings and in combination with the embodiments. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0113] In order to solve the technical problem of how to more timely and / or accurately determine whether the event concerned by the user occurs in the prior art, the present application provides an information pushing method, which can improve the accuracy and / or timeliness of the user determining whether the event concerned by the user occurs.
[0114] Figure 1A flowchart of an information pushing method provided in an embodiment of the present application is shown. The method can be applied to one or more electronic devices such as a smart phone, a notebook computer, a desktop computer, a portable computer, and a server. In addition, the execution subject of the method can be hardware or software. When the execution subject is hardware, the execution subject can be one or more of the electronic devices. For example, a single electronic device can execute the method, or multiple electronic devices can execute the method in cooperation with each other. When the execution subject is software, the method can be implemented as one or more software or software modules. No specific limitation is imposed herein.
[0115] As shown in Figure 1 , the method specifically includes the following steps.
[0116] In step 101, natural language and a video generated by camera monitoring are acquired, wherein the natural language is used to determine a video to be pushed.
[0117] In this embodiment, the natural language can be a language that evolves naturally with culture. For example, the language can be Chinese, English, or the like.
[0118] In some cases, the natural language can be collected via a terminal. When the execution subject of the method is a server, the terminal can send the natural language to the execution subject of the method after collecting the natural language. The terminal can be in communication connection with the execution subject of the method. When the execution subject of the method is a terminal, the natural language can be directly collected by the execution subject of the method.
[0119] The terminal can be hardware or software. For example, the terminal can be an electronic device such as a mobile phone or a computer, or an application running on the electronic device.
[0120] The natural language can be represented in the form of text or audio. For example, the natural language can be audio or text such as "remind me when the child comes home from tomorrow" or "remind me when the cat wakes up tomorrow".
[0121] The camera can be used to monitor a preset area, thereby generating a video of the preset area, i.e., a video generated by camera monitoring. The preset area can be the monitoring area of the camera.
[0122] In practice, the subject performing the method can first acquire the natural language and then acquire the video generated by the camera monitoring; or first acquire the video generated by the camera monitoring and then acquire the natural language; or simultaneously acquire the natural language and the video generated by the camera monitoring. The embodiment does not limit the acquisition order of the natural language and the video generated by the camera monitoring.
[0123] In step 102, the feature data of the natural language is determined to obtain first feature data.
[0124] In the embodiment, the first feature data can be the feature data of the natural language. As an example, the first feature data can be the semantic feature of the natural language.
[0125] In step 103, a first video is determined based on the video generated by the camera monitoring, and it is determined whether the first video matches the natural language based on at least two feature data of the first video and the first feature data.
[0126] In the embodiment, the first video can be any video directly collected by the camera, or a video generated after processing the video collected by the camera.
[0127] The at least two feature data of the first video can be data features obtained by performing feature extraction on the first video in at least two different data feature extraction manners.
[0128] In practice, the determination of whether the first video matches the natural language based on the at least two feature data of the first video and the first feature data can be performed in various manners.
[0129] As an example, the first video and the natural language can be input into a pre-trained artificial intelligence model, and the at least two feature data of the first video and the first feature data of the natural language are extracted via the artificial intelligence model, so as to obtain discrimination information indicating whether the first video matches the natural language.
[0130] The artificial intelligence model can be an artificial intelligence model trained in a supervised manner based on training samples including sample videos, sample natural languages, and sample discrimination information. The sample discrimination information indicates whether the sample video matches the sample natural language.
[0131] In addition, the determination of whether the first video matches the natural language can also be performed in other manners, which will be described below and will not be described here in detail.
[0132] Step 104, in the case that the first video matches the natural language, taking the first video as a to-be-pushed video, and pushing video information of the to-be-pushed video, wherein the video information represents information of the to-be-pushed video.
[0133] In the embodiment, in the case that the execution subject of the method is a terminal, the execution subject can push the video information of the to-be-pushed video to a user or the like by displaying the video information of the to-be-pushed video. For example, the execution subject can be a smart phone, a computer or the like.
[0134] In the case that the execution subject of the method is a server, the execution subject can send the video information of the to-be-pushed video to a terminal (for example, a smart phone, a computer) used by a user, so as to push the video information of the to-be-pushed video to the terminal. After receiving the video information of the to-be-pushed video, the terminal can display the video information of the to-be-pushed video.
[0135] The to-be-pushed video can be used to push to the terminal or the user or the like.
[0136] The video information can be prompt information, an address or the like of the to-be-pushed video, or can be the to-be-pushed video itself.
[0137] In practice, in the case that the first video matches the natural language, the first video can be pushed as a to-be-pushed video, or the video information (for example, prompt information) of the to-be-pushed video can be pushed first, and in the case that a preset operation (for example, a video playing operation) is detected through the video information, the first video can be pushed as a to-be-pushed video.
[0138] In some optional implementation manners of the embodiment, the first video can be generated in the following manner:
[0139] Firstly, an event frame is extracted from a video generated by a camera monitor.
[0140] The event frame can be one or more video frames representing an event.
[0141] The event frame can be extracted from the video generated by the camera monitor in various manners, which will be described below, and will not be described here.
[0142] Secondly, a plurality of event frames representing the same event are determined as the first video.
[0143] The similarity between the plurality of event frames representing the same event is greater than or equal to a preset first threshold, and the similarity between the plurality of event frames representing different events is less than the preset first threshold.
[0144] Here, the obtained video can include video frames of multiple different events. Thus, multiple event frames corresponding to each event can be determined as a first video. That is, each event can correspond to a first video. For example, the obtained video can include video frames representing event 1 as follows: video frame 5, video frame 6, and video frame 7, and also include video frames representing event 2 as follows: video frame 15, video frame 16, video frame 17, and video frame 19. Thus, a video composed of video frame 5, video frame 6, and video frame 7 can be determined as a first video, and a video composed of video frame 15, video frame 16, video frame 17, and video frame 19 can be determined as another first video.
[0145] It can be understood that in the optional implementation described above, the first video can be extracted from the obtained video continuously captured by the camera in units of events. Thus, once it is determined that the event triggered by the natural language expression is in the video captured by the camera, the video information of the event video can be pushed. Thus, the timeliness of the video information of the event triggered by the user's concern that the user receives can be further improved.
[0146] In some application scenarios in the optional implementation described above, the event frames can be extracted from the video generated by the camera monitoring in the following manner:
[0147] The event extraction model is used to extract the event frames from the video generated by the camera monitoring.
[0148] On this basis, the event extraction model can be trained in the following manner:
[0149] First, a training sample set is obtained.
[0150] The training sample in the training sample set includes a video (that is, a sample video), an event time (a sample event time), and an event label (that is, a sample event label).
[0151] The event time can represent the start time and the end time of the event, or the event time can represent the position of the event frame corresponding to the event in the video.
[0152] The event label can be the name of the event, for example, the event label can be "electric shock" or "approaching a safe".
[0153] Third, a machine learning algorithm is used to train the event extraction model by taking the video included in the training sample in the training sample set as input data and taking the event time and the event label as expected output data.
[0154] It can be understood that in the above application scenarios, the video, event time and event label can be used for training of the event extraction model, and then the event extraction model is used to extract the event frame in the video. In this way, the auxiliary extraction of the event frame through the event label can improve the accuracy of the event frame extraction.
[0155] In some application scenarios of the above optional implementation, the following steps can also be performed:
[0156] In a first step, the playing speed of the non-target video segment in the acquired video is determined as a first speed.
[0157] In a second step, the playing speed of the target video segment in the acquired video is determined as a second speed.
[0158] The target video segment is composed of event frames, in other words, the target video segment is an event video.
[0159] The second speed is less than the first speed.
[0160] It can be understood that in the above application scenarios, the target video segment can be played at a speed lower than that of the non-target video segment, which can enhance the immersive experience of the user and reduce the cost of recording life.
[0161] In some application scenarios of the above optional implementation, the following step can also be performed: generating a description text of the first video.
[0162] The description text can be used to describe the content of the first video.
[0163] In practice, an artificial intelligence model (such as a large language model) can be used to generate the description text of the first video.
[0164] As an example, a first video composed of a single frame or multiple frames is input into an artificial intelligence model, and the number of characters of the output text is set. The artificial intelligence model can output a description text for the first video. For example, a first video of a child returning home with his mother is input into an artificial intelligence model, and the artificial intelligence model is set to output a description of no more than 30 characters. The description includes time, location, characters and their attributes, behavior, etc. The artificial intelligence model can output "Today 11:30 at the door, a small boy wearing blue clothes returns home with his mother".
[0165] In some cases, the description text is used for terminal display. For example, in the case where the above execution subject is a server, the execution subject can send the description text to the terminal to make the terminal display the description text. In the case where the above execution subject is a terminal, the execution subject can directly display the description text.
[0166] As an example, please refer to Figure 3 In the above-mentioned application scenarios, the terminal displays the description texts "08:30, Robert and Lisa, Robert and Lisa go home with a skateboard", "07:10, express, the express delivery man wearing a blue hat delivers a package to the home and then immediately leaves". Figure 3
[0167] It can be understood that in the above-mentioned application scenarios, the user of the terminal can obtain the content of the first video through the description text without playing the video, so that the user can obtain the content of the first video more quickly.
[0168] In some application scenarios of the above-mentioned optional implementation manners, the following step can also be performed: generating the description text of the first video.
[0169] Here, the above-mentioned steps in the application scenarios can be implemented by referring to the above-mentioned application scenarios, please refer to the above description, and here is not described in detail.
[0170] In some cases, the description text is used to determine whether the first video is a search result of a video search request sent by the terminal. For example, in the case where the above-mentioned execution subject is a server, the execution subject can receive the video search request sent by the terminal, and determine whether the first video is a search result of the video search request based on the description text. For another example, in the case where the above-mentioned execution subject is a terminal, the execution subject can obtain a video search request input by a user or the like, and determine whether the first video is a search result of the video search request based on the description text.
[0171] Wherein, the video search request is used to perform video search.
[0172] In practice, whether the first video is a search result of the video search request can be determined by calculating the similarity between the description text and the first video.
[0173] It can be understood that in the above-mentioned application scenarios, the description text of the first video can be used to realize faster video search.
[0174] In some optional implementation manners of the present embodiment, the following steps can also be performed:
[0175] Firstly, determine at least one of the text and the music matched with the first video, to obtain the matching information of the first video.
[0176] Wherein, the matching information can include at least one of the text and the music matched with the first video.
[0177] Secondly, the matching information is fused with the first video to obtain a second video.
[0178] The second video can be a result of fusing the matching information with the first video. For example, the second video can be a video obtained after adding the matching text to the first video, or the second video can be a video obtained after adding the matching music to the first video.
[0179] In practice, the association between the text and / or music corresponding to each type of first video can be established first. Thus, the matching information of the first video can be determined by determining the text and / or music associated with the first video.
[0180] Thirdly, when a target operation for the second video is detected, the target operation is performed on the second video.
[0181] The target operation includes at least one of sharing, downloading, storing, and sending.
[0182] It can be understood that in the above optional implementation, the matching information of the first video can be determined automatically, and then the second video can be generated by fusing the first video and the matching information, so as to perform the sharing, downloading, storing, and sending of the second video more quickly.
[0183] In some optional implementations of the embodiment, the following steps can also be performed:
[0184] Firstly, one or more target video frames are determined from the first video.
[0185] The similarity between the target video frame and the preceding video frame is less than or equal to a preset second threshold, and the similarity between the target video frame and the following video frame is less than or equal to the preset second threshold. The preceding video frame is the previous video frame of the target video frame in the first video. The following video frame is the next video frame of the target video frame in the first video.
[0186] The preset second threshold can be equal to or different from the above-mentioned preset first threshold. In some cases, the preset second threshold can be greater than the first similarity threshold, so that the highlight video frame can be determined more accurately.
[0187] In some cases, a machine learning model can be used to determine the target video frame from the first video.
[0188] The machine learning model can use unsupervised contrastive learning, measure the similarity between frames in the image encoder (e.g., the similarity of images and audio over time), and define video frames with large differences as target video frames, i.e., highlight video frames.
[0189] Second, the target video frame is determined as the highlight video frame in the first video.
[0190] In some cases, the highlight video frame can be one or more consecutive video frames in the first video. On this basis, for the video frames in the first video located before the highlight video frame, the number of video frames between them (located before the highlight video frame) and the highlight video frame is positively correlated with their (located before the highlight video frame) playback speed, i.e., the more video frames between them and the highlight video frame, the faster their playback speed; for the video frames in the first video located after the highlight video frame, the number of video frames between them (located after the highlight video frame) and the highlight video frame is negatively correlated with their (located before the highlight video frame) playback speed, i.e., the more video frames between them and the highlight video frame, the slower their playback speed.
[0191] It can be understood that in the above optional implementation, the highlight video frame in the first video can be determined.
[0192] In some optional implementations of the present embodiment, the method is applied to the first device end.
[0193] Among them, the first device end can represent a terminal, or a server. As an example, the first device end can be a camera.
[0194] On this basis, the video information of the to-be-pushed video can be pushed in the following way:
[0195] First, obtain the location information of the second device end.
[0196] Among them, the second device end can be another device end different from the first device end. As an example, the second device end can represent another terminal or server different from the first device end. As an example, in the case where the first device end is a camera, the second device end can be a smartphone, a computer, etc.
[0197] The above location information can represent the location of the second device end.
[0198] Second, determine the video information of the to-be-pushed video based on the location information.
[0199] Here, after a correspondence between the pre-established location information and the video information is established, the video information corresponding to the location information obtained in the first step can be determined as the video information of the video to be pushed.
[0200] For example, the location information 1 can correspond to the video information 1 of the video to be pushed, and the location information 2 can correspond to the video information 2 of the video to be pushed.
[0201] Thirdly, the video information is pushed to the second device.
[0202] It can be understood that in the above implementation, in the case that the second device is located at different locations, different video information of the video to be pushed can be pushed to the second device.
[0203] In some application scenarios of the above optional implementation, the video information of the video to be pushed can be determined based on the location information in the following manner:
[0204] Firstly, the location of the camera is determined to obtain a target location.
[0205] The target location can represent the location of the camera.
[0206] Then, the distance between the location represented by the location information and the target location is determined to obtain a target distance.
[0207] The target distance can represent the distance between the location represented by the location information and the target location.
[0208] Then, it is determined whether the target distance is greater than or equal to a preset distance threshold.
[0209] Subsequently, in the case that the target distance is greater than or equal to the preset distance threshold, the video information of the video to be pushed is determined as first information. In the case that the target distance is less than the preset distance threshold, the video information of the video to be pushed is determined as second information.
[0210] The first information represents a request to control the camera to monitor a target region. The second information represents the location of the target region.
[0211] The target region is a region for which the camera generates the video to be pushed.
[0212] In the case that the second device is a smart phone, if the target distance is greater than or equal to the preset distance threshold, it can be considered that the user using the second device is not at home at this time; if the target distance is less than the preset distance threshold, it can be considered that the user using the second device is at home at this time.
[0213] It can be understood that in the above application scenario, when the distance between the camera and the second device end is far (for example, not at home), the user can be requested to control the camera to collect video of the event triggering area so that the user can monitor remotely. When the distance between the camera and the second device end is close (for example, at home), the location of the event triggering area can be informed to the user through the second device end so that the user can arrive at the event triggering area immediately.
[0214] The information push method provided in the embodiment of the present application can obtain natural language and video generated by camera monitoring, wherein the natural language is used to determine the video to be pushed, and then the feature data of the natural language is determined to obtain the first feature data, and then the first video is determined based on the video generated by the camera monitoring, and based on at least two feature data of the first video and the first feature data, it is determined whether the first video matches the natural language, and finally, if the first video matches the natural language, the first video is used as the video to be pushed, and the video information of the video to be pushed is pushed, wherein the video information represents the information of the video to be pushed. Thus, by setting the natural language in advance, an automatic reminder can be realized when an event represented by the natural language is triggered in the video generated by the camera monitoring, and the video information triggered by the corresponding event can be pushed more accurately and / or timely, thereby improving the accuracy and / or timeliness of users and other objects in judging whether the event they are concerned about has occurred.
[0215] Figure 2 This is a flow chart of another information push method provided in the embodiment of the present application. Figure 2 As shown, the method specifically includes:
[0216] Step 201: Obtain natural language and video generated by camera monitoring, wherein the natural language is used to determine the video to be pushed.
[0217] In this embodiment, step 201 and Figure 1 Step 101 in the corresponding embodiment is basically the same and will not be described again here.
[0218] Step 202: Determine the feature data of the natural language to obtain first feature data.
[0219] In this embodiment, the first feature data may be feature data of the natural language. As an example, the first feature data may be a semantic feature of the natural language.
[0220] Step 203: determining a first video based on the video generated by the camera monitoring, determining at least two feature data of the first video among image feature data, text feature data, and audio feature data, and obtaining second feature data.
[0221] In the embodiment, the second feature data can be at least two of the image feature data, the text feature data and the audio feature data of the first video.
[0222] In some cases, the second feature data can include the image feature data, the text feature data and the audio feature data of the first video.
[0223] In step 204, it is determined whether the first video matches the natural language based on the first feature data and the second feature data.
[0224] In the embodiment, the similarity between the first feature data and the second feature data can be calculated, and thus, it can be determined whether the first video matches the natural language based on the size relationship between the similarity and a preset similarity threshold. For example, if the similarity is greater than or equal to the preset similarity threshold, it can be determined that the first video matches the natural language. If the similarity is less than the preset similarity threshold, it can be determined that the first video does not match the natural language.
[0225] In step 205, in the case where the first video matches the natural language, the first video is taken as a to-be-pushed video, and video information of the to-be-pushed video is pushed, where the video information represents information of the to-be-pushed video.
[0226] In the embodiment, step 205 is basically the same as step 104 in the corresponding embodiment, which will not be described here. Figure 1
[0227] It should be noted that, in addition to the above-described content, the embodiment can also include the corresponding technical features described in the Figure 1 corresponding embodiments, thereby achieving the technical effects of the information pushing method shown in the Figure 1 embodiments. For details, please refer to the related description, which will not be described here for brevity. Figure 1
[0228] The information pushing method provided by the embodiment can determine whether the first video matches the natural language through the feature data of the natural language and the multi-modal fusion feature data of at least two dimensions of the image feature data, the text feature data and the audio feature data of the first video, so that the accuracy of determining whether the first video matches the natural language can be improved.
[0229] The embodiments of the present application will be described below, but it should be noted that the embodiments of the present application can have the features described below, but the following description does not constitute a limitation on the scope of protection of the embodiments of the present application.
[0230] Before introducing the present scheme, first introduce the technical terms involved in the present scheme as follows:
[0231] Multi-modal: refers to the form of data, such as text, audio, image, video, etc.
[0232] Key frame: refers to the frame where the key action of the character or object motion change is located.
[0233] Transformer: a sequence-based deep learning model technology.
[0234] Attention mechanism: the attention mechanism in deep learning is a method that simulates the human visual and cognitive system, which allows the neural network to focus on the relevant part when processing input data.
[0235] Text vector: the feature vector of text mapped to high-dimensional space by deep learning model.
[0236] Highlight moment: the most exciting moment in a video.
[0237] At present, the user's demand for family diary mainly includes four aspects:
[0238] The first aspect is to show the family diary in the form of text, helping users quickly analyze the main events of the day in the family when they are busy.
[0239] The second aspect is to show the family diary in the form of video, helping users to enjoy / review key events in an immersive state when they are at leisure.
[0240] The third aspect is to realize "event triggered reminder" by setting labels in advance, ensuring that users can respond to important family events in a timely manner.
[0241] The fourth aspect is to further generate high-quality short videos with pictures, texts, and music based on text and video diaries, capture the highlight moments, and play them at an adaptive speed to create an immersive experience and record life at a low cost.
[0242] In addition, there is currently no related scheme to show family security diary in the form of text. Users need to quickly understand the main family events in a day through text in some situations (such as outdoor busy), and pure "video highlights" family diary is not concise and clear enough, lacking a refined description of events. The video highlights diary form is not convenient for later retrieval and evidence collection. The generated video lacks immersive experience, mainly in that it cannot flexibly match the appropriate background music and cannot realize slow playback of highlight moments.
[0243] Therefore, the present method can solve the above technical problems through the functions provided by the following modules:
[0244] Natural language instruction setting module: the user inputs "remind me when the child comes home tomorrow", and the system converts this text into a text vector (i.e., the first feature data described above).
[0245] Event key frame extraction module: this module first uses a large model (i.e., a large language model, corresponding to the event extraction model described above) to generate a template for large-scale multi-event video data as "start time-end time-event" description text. Among them, the start time and end time correspond to the event time described above. The time corresponds to the event label described above. Then the generated data is used to train the large model to enhance the model's ability to perceive the time boundaries of different events in the video, so as to more accurately extract the start frame and end frame of the event.
[0246] Voice / text / image understanding module:
[0247] (1) Input single frame or multi-frame key frame (i.e., event frame described above) into image understanding module (multimodal model), and set the number of output text words. This module will output a description of the image or image set. For example, input a video of a child and mother coming home into the image understanding module and set the image understanding module to output a description of no more than 30 words. The description includes time, place, character and attribute, behavior, etc. information, then the module will output "today 11:30 at the door there is a small boy wearing blue clothes and mother come home together".
[0248] (2) When it is judged that the event has ended, the image or video is converted into a picture vector and matched with the text vector of the natural language instruction setting module. When the similarity is higher than the specified threshold, the APP (Application, application program) push message (i.e., the video information described above) is triggered to remind.
[0249] Diary automatic editing and highlight moment capture module: this module splices all related events of the day in chronological order, and automatically understands the video content based on the multimodal model to generate background music, scripts for videos in different time periods, and capture highlight moments, and adaptively play speed.
[0250] (1) The module first uses a text encoder, a speech encoder, and an image encoder to extract the text, audio, and image features in the video respectively, corresponding to the second feature data mentioned above, and uses a transformer-based multimodal fusion network to generate a multimodal representation that integrates text, audio, and images. In order to better understand and fuse multimodal information, a squeeze-excitation attention mechanism is introduced to calculate the correlation between the same modality, different channels, and different modal features, thereby enhancing the model's attention to related features and achieving feature enhancement. Finally, based on the fused features, the video content is understood, background music that matches the theme is matched, and a descriptive text is generated. In addition, by using self-supervised contrastive learning to measure the similarity of images between frames in the image encoder and the audio similarity in the time series, the method defines videos with large differences between the previous and the next as highlights / key moments (i.e., the above-mentioned highlight video frames), and adaptively reduces the playback speed of these moments, effectively improving the immersion of the family diary.
[0251] Sharing method: The edited short video can be directly shared to short video social software.
[0252] Typical application scenarios are as follows:
[0253] Scene 1: The child returns home:
[0254] The user enters text or voice in the APP's message push box, "Remind me when the child comes home from tomorrow" (the natural language mentioned above). The system converts this text into a text vector (the first feature data mentioned above). Starting from the second day of setting the instruction, the recorded pictures and videos will be segmented by events, and different events will be understood and descriptive text will be generated. When the similarity between the event vector and the text vector is higher than the set threshold, the APP pushes a message (that is, a video message) to remind you. At 24:00 every day, the number of events that occurred that day is counted, and the videos of each event are edited and the corresponding text descriptions are generated to form a multimedia family diary, which is convenient for users to save and forward and share.
[0255] Scene 2, pet activities:
[0256] The user inputs text or voice "Remind me when the cat wakes up tomorrow" (i.e., the natural language described above) in the message push box of the APP. The system converts this text into a text vector (i.e., the first feature data described above). Starting from the second day of setting the instruction, the video recorded by the terminal device will be segmented by events as the dimension, and different events will be understood and text descriptions will be generated. When the event vector of the video and the text vector described in the push instruction have a similarity higher than a set threshold, the APP pushes a message (i.e., the video information described above) for reminding. At 24:00, the number of events occurring on the same day is counted, and all the videos of the cat's activities after waking up are spliced in chronological order, and the corresponding text description is generated, such as "9:00, the cat starts to move" (normal playback speed); "12:00, the cat wakes up again and starts to drink water" (normal playback speed); "15:00, the cat jumps off the sofa" (3 / 5 normal playback speed), and the video is matched with light and warm background music. It is worth noting that the video of "15:00, the cat jumps off the sofa" is determined by the system as the highlight moment of the cat (corresponding to the highlight video frame described above), so the playback speed of this video segment is switched to 3 / 5 of the normal speed, and finally an immersive multimedia family diary is formed, which is convenient for the user to save, forward and share.
[0257] It should be noted that, in addition to the above described content, the present embodiment can also include the technical features described in the above embodiments, thereby achieving the technical effects of the information push method shown above. For details, please refer to the above description. For the sake of brevity, no further description is given here.
[0258] The information push method provided by the present embodiment utilizes a multi-modal large model to convert natural language into a message push instruction, which can flexibly set a prompt task and adopt a text vector and picture / video vector comparison method as a triggering object. The use of a large model enhances the time boundary perception ability of different events occurring and ending in a video, which can more accurately extract the key frames (i.e., the event frames described above) of related events. The use of a multi-modal model summarizes the video of a home security monitoring camera into text and clips the video into a family highlight according to the event dimension to generate a multimedia family diary in the form of graphs, texts and sounds. By capturing the highlight moment to adaptively adjust the playback speed of the diary, the user's immersive experience is enhanced, and the cost of recording life is reduced. The method can record the events occurring in the family every day in multiple dimensions of text, pictures and videos, which has the characteristics of concise and clear text diary and convenient retrieval, and also enhances the characteristics of intuitive and vividness and emotional communication. The more powerful video understanding technology realizes the fusion of three modal data and automatically identifies the highlight moment in the video to adaptively adjust the playback speed, thereby creating an immersive user experience.
[0259] Figure 4 A structural schematic diagram of an information push device provided by the present embodiment. The device is arranged on a server and specifically comprises:
[0260] The acquisition unit 401 is configured to acquire natural language and a video generated by camera monitoring, wherein the natural language is used to determine a to-be-pushed video.
[0261] The first determination unit 402 is configured to determine feature data of the natural language to obtain first feature data.
[0262] The second determination unit 403 is configured to determine a first video based on the video generated by the camera monitoring, and determine whether the first video matches the natural language based on at least two kinds of feature data of the first video and the first feature data.
[0263] The pushing unit 404 is configured to, in a case where the first video matches the natural language, push the first video as a to-be-pushed video, and push video information of the to-be-pushed video, wherein the video information represents information of the to-be-pushed video.
[0264] In one possible implementation, the determination of whether the first video matches the natural language comprises:
[0265] The determination of the feature data of the natural language to obtain the first feature data.
[0266] The determination of at least two kinds of feature data of image feature data, text feature data and audio feature data of the first video to obtain second feature data.
[0267] The determination of whether the first video matches the natural language based on the first feature data and the second feature data.
[0268] In one possible implementation, the first video is generated in the following manner:
[0269] Extracting event frames from a video generated by camera monitoring.
[0270] Determining, as the first video, multiple event frames representing a same event, wherein a similarity between the multiple event frames representing the same event is greater than or equal to a preset first threshold, and a similarity between multiple event frames representing different events is less than the preset first threshold.
[0271] In one possible implementation, the extraction of the event frames from the video generated by the camera monitoring comprises:
[0272] Extracting the event frames from the video generated by the camera monitoring by using an event extraction model; and
[0273] The event extraction model is trained in the following manner:
[0274] A training sample set is obtained, wherein a training sample in the training sample set comprises a video, an event time, and an event label;
[0275] A machine learning algorithm is used to train an event extraction model, wherein the video included in the training sample in the training sample set is used as input data, and the event time and the event label are used as expected output data.
[0276] In one possible implementation, the apparatus further comprises:
[0277] A third determination unit (not shown in the figure) is configured to determine a first speed as a playing speed of a non-target video segment in the obtained video;
[0278] A fourth determination unit (not shown in the figure) is configured to determine a second speed as a playing speed of a target video segment in the obtained video;
[0279] The target video segment is composed of event frames, and the second speed is less than the first speed.
[0280] In one possible implementation, the apparatus further comprises:
[0281] A generation unit (not shown in the figure) is configured to generate a description text of the first video;
[0282] The description text is used for terminal display, and / or the description text is used to determine whether the first video is a search result of a video search request sent by the terminal.
[0283] In one possible implementation, the apparatus further comprises:
[0284] A fifth determination unit (not shown in the figure) is configured to determine at least one of a text and music matched with the first video, to obtain matching information of the first video;
[0285] A fusion unit (not shown in the figure) is configured to fuse the matching information with the first video to obtain a second video;
[0286] A second processing unit (not shown in the figure) is configured to perform a target operation on the second video in a case where the target operation for the second video is detected, wherein the target operation comprises at least one of sharing, downloading, storing, and sending.
[0287] In one possible implementation, the apparatus further comprises:
[0288] A sixth determining unit (not shown in the figure) is configured to determine a target video frame from the first video, wherein a similarity between the target video frame and a previous video frame is less than or equal to a preset second threshold, and a similarity between the target video frame and a subsequent video frame is less than or equal to the preset second threshold, the previous video frame being a previous frame of the target video frame in the first video, and the subsequent video frame being a subsequent frame of the target video frame in the first video.
[0289] A seventh determining unit (not shown in the figure) is configured to determine the target video frame as a highlight video frame in the first video.
[0290] In one possible implementation, the method is applied to a first device end; and
[0291] The video information of the to-be-pushed video includes:
[0292] Obtaining position information of a second device end;
[0293] Determining the video information of the to-be-pushed video based on the position information;
[0294] Pushing the video information to the second device end.
[0295] In one possible implementation, the determining of the video information of the to-be-pushed video based on the position information includes:
[0296] Determining a position where the camera is located to obtain a target position;
[0297] Determining a distance between the position represented by the position information and the target position to obtain a target distance;
[0298] Determining whether the target distance is greater than or equal to a preset distance threshold;
[0299] In a case where the target distance is greater than or equal to the preset distance threshold, determining that the video information of the to-be-pushed video is first information; wherein the first information indicates a request for controlling the camera to monitor a target area, and the target area is an area from which the to-be-pushed video is generated;
[0300] In a case where the target distance is less than the preset distance threshold, determining that the video information of the to-be-pushed video represents second information; wherein the second information indicates a position of a target area, and the target area is an area from which the to-be-pushed video is generated.
[0301] The information pushing apparatus provided in this embodiment can be, for example, Figure 4The information pushing device shown in the embodiment can perform all steps of the information pushing method applied to the server side, and further realize the technical effects of the information pushing method applied to the server side. For brevity, the related description is not repeated here.
[0302] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in the figure, Figure 5 The electronic device 500 shown in the figure includes at least one processor 501, a memory 502, at least one network interface 504 and other user interfaces 503. The various components in the electronic device 500 are coupled together through a bus system 505. It can be understood that the bus system 505 is used to realize the connection communication between the components. In addition to including a data bus, the bus system 505 also includes a power bus, a control bus and a status signal bus. However, for the sake of clear illustration, all kinds of buses are marked as the bus system 505 in the figure. Figure 5
[0303] The user interface 503 can include a display, a keyboard or a clicking device (for example, a mouse, a trackball, a touchpad or a touch screen, etc.).
[0304] It is to be understood that the memory 502 in embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, without being limited to, these and any other suitable types of memory.
[0305] In some embodiments, the memory 502 stores the following elements, executable units or data structures, or a subset of them, or an extended set of them: an operating system 5021 and an application program 5022.
[0306] Among them, the operating system 5021 contains various system programs, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 5022 contains various application programs, such as Media Player, Browser, etc., for implementing various application services. The program for implementing the method of the embodiments of the present application can be contained in the application program 5022.
[0307] In the present embodiment, by calling the program or instruction stored in the memory 502, specifically, the program or instruction stored in the application program 5022, the processor 501 is used to execute the method steps provided by each method embodiment, for example, including:
[0308] Acquire natural language and video generated by camera monitoring, wherein the natural language is used to determine a video to be pushed;
[0309] Determine feature data of the natural language to obtain first feature data;
[0310] Determine a first video based on the video generated by the camera monitoring, and determine whether the first video matches the natural language based on at least two feature data of the first video and the first feature data;
[0311] In a case where the first video matches the natural language, take the first video as a video to be pushed, and push video information of the video to be pushed, wherein the video information represents information of the video to be pushed.
[0312] The method disclosed in the embodiments of the present application can be applied to the processor 501 or implemented by the processor 501. The processor 501 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 501. The processor 501 mentioned above can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software units in the code processor for execution. The software unit can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage media in the art. The storage medium is located in the memory 502, and the processor 501 reads the information in the memory 502 and combines the hardware to complete the steps of the above method.
[0313] It can be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing units can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP Devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described above, or a combination thereof.
[0314] For software implementation, the techniques described herein can be implemented with a processing unit executing program code embodied in software. The software is stored in a storage and executed by the processing unit. The storage can be implemented within the processing unit or external to the processing unit.
[0315] The electronic device provided by the embodiments can be an electronic device as shown in Figure 5 The electronic device provided by the embodiments can be an electronic device as shown in
[0316] The embodiments of the present application further provide a storage medium (computer readable storage medium). The storage medium stores one or more programs. The storage medium can include a volatile memory, such as a random access memory, and can also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk, or a solid state disk. The storage medium can also include a combination of the above-mentioned memories.
[0317] When the one or more programs stored in the storage medium are executed by the one or more processors, the information pushing method executed at the electronic device side described above can be implemented.
[0318] The processor is configured to execute the information pushing program stored in the memory to implement the steps of the information pushing method executed at the electronic device side described above.
[0319] Obtaining natural language and a video generated by a camera monitoring, wherein the natural language is used to determine a video to be pushed;
[0320] Determining feature data of the natural language to obtain first feature data;
[0321] determine a first video based on the video generated by the camera monitoring, and determine whether the first video matches the natural language based on at least two feature data of the first video and the first feature data;
[0322] in a case where the first video matches the natural language, take the first video as a to-be-pushed video, and push video information of the to-be-pushed video, wherein the video information represents information of the to-be-pushed video.
[0323] Those skilled in the art should further appreciate that the units and algorithm steps of various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of various examples have been described in general terms above. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0324] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, software executed by a processor, or a combination of both. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.
[0325] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has", "having" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be interpreted as necessarily requiring their performance in the specific order indicated, unless explicitly specified otherwise. It should also be understood that additional or alternative steps can be employed.
[0326] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and it is intended to embrace all such modifications and changes that fall within the scope of the application. Accordingly, the application is not to be restricted in scope to the specific embodiments disclosed herein but is to be accorded the full scope that the principles and novel features request appropriately granted.
Claims
1. An information push method characterized by comprising: The method comprises: acquiring natural language and a video generated by camera monitoring, wherein the natural language is used to determine a video to be pushed; determining feature data of the natural language to obtain first feature data; determining a first video based on the video generated by the camera monitoring, and determining whether the first video matches the natural language based on at least two feature data of the first video and the first feature data; in a case where the first video matches the natural language, taking the first video as a video to be pushed, and pushing video information of the video to be pushed, wherein the video information represents information of the video to be pushed.
2. The method of claim 1, wherein, The determination of whether the first video matches the first feature data based on the at least two feature data of the first video comprises: determining at least two feature data of image feature data, text feature data and audio feature data of the first video to obtain second feature data; determining whether the first video matches the natural language based on the first feature data and the second feature data.
3. The method of claim 1, wherein, The first video is generated in the following manner: extracting event frames from a video generated by camera monitoring; determining multiple event frames representing the same event as the first video, wherein a similarity between the multiple event frames representing the same event is greater than or equal to a preset first threshold, and a similarity between multiple event frames representing different events is less than the preset first threshold.
4. The method of claim 3, wherein, The extraction of the event frames from the video generated by the camera monitoring comprises: extracting the event frames from the video generated by the camera monitoring by using an event extraction model; and The event extraction model is trained in the following manner: acquiring a training sample set, wherein a training sample in the training sample set comprises a video, an event time and an event label; training an event extraction model by using a machine learning algorithm, taking the video included in the training sample in the training sample set as input data, and taking the event time and the event label as expected output data.
5. The method of claim 3, wherein, The method further comprises: determining a playing speed of a non-target video segment in the acquired video as a first speed; determining a playing speed of a target video segment in the acquired video as a second speed; wherein the target video segment is composed of event frames, and the second speed is less than the first speed.
6. The method of claim 3, wherein, The method further comprises: generating a description text of the first video; wherein the description text is used for terminal display, and / or the description text is used to determine whether the first video is a search result of a video search request sent by a terminal.
7. The method according to one of claims 1 to 6, characterized in that The method further comprises: determining at least one of text and music matched with the first video to obtain matching information of the first video; fusing the matching information with the first video to obtain a second video; in a case where a target operation for the second video is detected, performing the target operation on the second video, wherein the target operation comprises at least one of sharing, downloading, storing and sending.
8. The method according to one of claims 1 to 6, characterized in that The method further comprises: determining a target video frame from the first video, wherein a similarity between the target video frame and a previous video frame is less than or equal to a preset second threshold, and a similarity between the target video frame and a next video frame is less than or equal to the preset second threshold, the previous video frame being a previous frame of the target video frame in the first video, and the next video frame being a next frame of the target video frame in the first video; determining the target video frame as a highlight video frame in the first video.
9. The method according to one of claims 1 to 6, characterized in that The method is applied to a first device end; and The video information of the video to be pushed includes: obtaining position information of a second device end; determining the video information of the video to be pushed based on the position information; pushing the video information to the second device end.
10. The method of claim 9, wherein, The determination of the video information of the video to be pushed based on the position information includes: determining a position where the camera is located to obtain a target position; determining a distance between the position represented by the position information and the target position to obtain a target distance; determining whether the target distance is greater than or equal to a preset distance threshold; in a case where the target distance is greater than or equal to the preset distance threshold, determining that the video information of the video to be pushed is first information, wherein the first information represents a request for controlling the camera to monitor a target area, and the target area is an area for generating the video to be pushed; in a case where the target distance is less than the preset distance threshold, determining that the video information of the video to be pushed represents second information, wherein the second information represents a position of a target area, and the target area is an area for generating the video to be pushed.
11. An information push apparatus characterized by comprising: The apparatus includes: an obtaining unit configured to obtain a natural language and a video generated by camera monitoring, wherein the natural language is used to determine a video to be pushed; a first determining unit configured to determine feature data of the natural language to obtain first feature data; a second determining unit configured to determine a first video based on the video generated by the camera monitoring, and determine whether the first video matches the natural language based on at least two types of feature data of the first video and the first feature data; a pushing unit configured to, in a case where the first video matches the natural language, push the first video as a video to be pushed, and push video information of the video to be pushed, wherein the video information represents information of the video to be pushed.
12. An electronic device, comprising: including: a memory configured to store a computer program; a processor configured to execute the computer program stored in the memory, and when the computer program is executed, the information pushing method of any one of claims 1-10 is implemented.
Citation Information
Cited By
Method and system for generating monitoring video user attention information based on large model
CN121725409A