Video interaction method, device and equipment based on user requirements

By analyzing user needs to obtain location information of the area of interest, using video surveillance equipment for real-time monitoring and transmitting videos, the problem of unreliable video interaction information in the prior art is solved, and high-reliability real-time video interaction is achieved.

CN119815125BActive Publication Date: 2025-07-11CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411968955.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-11
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In the existing video interaction technology, the video information obtained by users lacks real-time and reliability, resulting in unreliable information.

Method used

By obtaining the interaction request information of the target user, analyzing the location information of the area of interest, using the video surveillance device deployed in the area for real-time video surveillance, and transmitting the monitoring video to the user terminal, real-time video interaction between the electronic device and the terminal device is realized.

Benefits of technology

It improves the real-time and reliability of video interactions, ensuring that the information obtained by users has high accuracy and timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815125B_ABST
    Figure CN119815125B_ABST
Patent Text Reader

Abstract

The video interaction method, device and equipment based on user requirements provided by this application relate to the technical field of video interaction. In this application, first, obtain the target interaction request information sent by the target user through the corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user; secondly, based on the target location information, determine the target real-time monitoring video; then, transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device. Based on the above content, the problem of relatively low reliability of video interaction existing in the prior art can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video interaction, and more particularly, to a video interaction method, apparatus, and device based on user requirements. Background Art

[0002] With the popularization of intelligent devices and the continuous progress of technology, video interaction technology has been widely applied in multiple fields, including telemedicine, video surveillance, online education, video conferencing, customer service, etc. Among them, video surveillance can be used as a basis for users' travel choices. For example, in the prior art, users can search for videos of areas of interest to understand the situation of corresponding areas. However, the videos searched by users are generally based on those sent by other users, so they generally do not have real-time nature, making the information obtained unreliable, that is, there is a problem of relatively low reliability in video interaction. Summary of the Invention

[0003] In view of this, the purpose of the present application is to provide a video interaction method, apparatus, and device based on user requirements to improve the problem of relatively low reliability in video interaction existing in the prior art.

[0004] To achieve the above object, the present application adopts the following technical solutions:

[0005] A video interaction method based on user requirements, which is applied to an electronic device. The video interaction method based on user requirements includes:

[0006] Obtain target interaction request information sent by a target user through a corresponding target terminal device, and parse the target interaction request information to obtain target location information corresponding to the target user, where the target location information is used to reflect the area of interest of the target user;

[0007] Based on the target location information, determine a target real-time surveillance video, where the target real-time surveillance video is obtained by a video surveillance device deployed in the area of interest reflected by the target location information performing real-time video surveillance on the area of interest;

[0008] Transmit the target real-time surveillance video to the target terminal device corresponding to the target user to complete real-time video interaction between the electronic device and the target terminal device.

[0009] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of determining a target real-time surveillance video based on the target location information includes:

[0010] Determine each video surveillance device deployed in the region of interest reflected by the target location information to form at least one target video surveillance device;

[0011] Control each of the target video surveillance devices to start and perform video surveillance respectively, and obtain the target real-time surveillance video corresponding to each of the target video surveillance devices.

[0012] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of determining each video surveillance device deployed in the region of interest reflected by the target location information to form at least one target video surveillance device includes:

[0013] Determine each video surveillance device deployed in the region of interest reflected by the target location information to form at least one candidate video surveillance device;

[0014] When the number of devices of the at least one candidate video surveillance device is less than or equal to a pre-determined reference number, use the at least one candidate video surveillance device as the target event surveillance device;

[0015] When the number of devices of the at least one candidate video surveillance device is greater than the reference number, based on a pre-determined device screening rule, determine at least one candidate video surveillance device from the at least one candidate video surveillance device as the target event surveillance device.

[0016] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of, when the number of devices of the at least one candidate video surveillance device is greater than the reference number, based on a pre-determined device screening rule, determining at least one candidate video surveillance device from the at least one candidate video surveillance device as the target event surveillance device includes:

[0017] When the number of devices of the at least one candidate video surveillance device is greater than the reference number, obtain the regional description information of the local surveillance area corresponding to each candidate video surveillance device, where the regional description information includes video and / or text;

[0018] Obtain the target user requirement information sent by the target user through the target terminal device, where the target user requirement information is used to reflect the content of interest of the target user;

[0019] For each candidate video surveillance device, determine the matching degree between the regional description information of the local surveillance area corresponding to the candidate video surveillance device and the target user requirement information;

[0020] Each candidate video surveillance device with a matching degree greater than or equal to a pre-determined reference matching degree is used as a target event surveillance device, or, the target number of candidate video surveillance devices with the largest matching degree is used as the target event surveillance device.

[0021] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of determining the target real-time surveillance video based on the target location information includes:

[0022] Determine the video surveillance devices deployed in the region of interest reflected by the target location information to form target video surveillance devices, where the video surveillance devices belong to the airborne devices of the unmanned aerial vehicle;

[0023] Control the target video surveillance devices to move in the region of interest reflected by the target location information, and perform video surveillance during the movement to obtain the corresponding target real-time surveillance video.

[0024] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of controlling the target video surveillance devices to move in the region of interest reflected by the target location information and performing video surveillance during the movement to obtain the corresponding target real-time surveillance video includes:

[0025] Obtain the target user demand information sent by the target user through the target terminal device, where the target user demand information is used to reflect the content of interest of the target user;

[0026] Based on the target user demand information, determine a target movement path in the region of interest reflected by the target location information, where the target movement path includes at least a starting point and an ending point;

[0027] Control the target video surveillance devices to move along the target movement path in the region of interest reflected by the target location information, and perform video surveillance through the target video surveillance devices during the movement to obtain the corresponding target real-time surveillance video.

[0028] In a preferred selection of the present application, in the above video interaction method based on user requirements, the step of determining a target movement path in the region of interest reflected by the target location information based on the target user demand information includes:

[0029] Based on the target user demand information, determine each historical user having demand relevance with the target user to form at least one relevant historical user, where the historical user refers to a user who has performed real-time video interaction with the electronic device in history;

[0030] Obtain the historical real-time monitoring videos formed by each of the relevant historical users in history, and form at least one historical real-time monitoring video;

[0031] Mine and aggregate multi-modal features of the at least one historical real-time monitoring video and the target user demand information to form the multi-modal user demand features corresponding to the target user, where the multi-modal user demand features are used to reflect the semantic information of the at least one historical real-time monitoring video and the target user demand information;

[0032] Perform path planning processing on the region of interest reflected by the target location information to form multiple candidate movement paths, and for each of the candidate movement paths, respectively obtain the node description information of each path node in the candidate movement path, and combine the node description information of each path node to form the path description information corresponding to the candidate movement path, where the node description information includes videos and / or texts;

[0033] Respectively perform feature mining on the path description information corresponding to each of the candidate movement paths to form the path description features corresponding to each of the candidate movement paths;

[0034] Respectively determine the similarity between the path description features corresponding to each of the candidate movement paths and the multi-modal user demand features, and determine the candidate movement path with the maximum similarity as the target movement path.

[0035] In a preferred selection of the present application, in the above video interaction method based on user needs, the step of determining each historical user having demand relevance with the target user based on the target user demand information to form at least one relevant historical user includes:

[0036] Determine the historical user relationship network constructed based on each historical user and the historical user demand information corresponding to each historical user, and perform a first update on the historical user relationship network based on the target user and the target user demand information to form a first updated user relationship network. During the first update, connect the historical users corresponding to each historical user demand information with a relevance greater than or equal to the preset relevance to the target user demand information and the target user to represent that the historical user has demand relevance with the target user;

[0037] Respectively perform feature mining on each of the to-be-processed user demand information and output the to-be-processed user demand features corresponding to each of the to-be-processed user demand information, where each of the to-be-processed user demand information belongs to the historical user demand information or the target user demand information;

[0038] For each to-be-processed user in the first updated user relationship network, in the first updated user relationship network, determine a local user relationship network centered on the to-be-processed user, where each to-be-processed user belongs to the historical user or the target user;

[0039] Perform random walks in the first updated user relationship network to form a plurality of random walk paths, and randomly serialize the plurality of random walk paths to form a corresponding target path sequence, where the path end point of each random walk path is the target user;

[0040] According to the sequence relationship of the plurality of random walk paths in the target path sequence, and according to the direction from the path start point to the path end point in the random walk path, determine the currently traversed to-be-processed user, and based on the current to-be-processed user demand characteristics corresponding to each to-be-processed user in the local user relationship network of the currently traversed to-be-processed user, perform associated mining on the current to-be-processed user demand characteristics corresponding to the currently traversed to-be-processed user to obtain corresponding associated user demand characteristics, and use the associated user demand characteristics as the new current to-be-processed user demand characteristics, and after performing associated mining on the current to-be-processed user demand characteristics corresponding to the target user in the last random walk path, based on the feature similarity between the current to-be-processed user demand characteristics corresponding to each to-be-processed user, perform a second update on the first updated user relationship network to form a second updated user relationship network;

[0041] In the second updated user relationship network, determine each historical user having a connection relationship with the target user to form at least one relevant historical user.

[0042] This application also provides a video interaction device based on user needs, which is applied to an electronic device. The video interaction device based on user needs includes:

[0043] A location information determination module, configured to obtain target interaction request information sent by a target user through a corresponding target terminal device, and parse the target interaction request information to obtain target location information corresponding to the target user, where the target location information is used to reflect the area of interest of the target user;

[0044] A monitoring video determination module, configured to determine a target real-time monitoring video based on the target location information, where the target real-time monitoring video is obtained by a video monitoring device deployed in the area of interest reflected by the target location information to perform real-time video monitoring on the area of interest;

[0045] The monitoring video transmission module is used to transmit the target real-time monitoring video to the target terminal device corresponding to the target user, so as to complete the real-time video interaction between the electronic device and the target terminal device.

[0046] On this basis, the present application further provides an electronic device, including:

[0047] A memory for storing a computer program;

[0048] A processor connected to the memory, for executing the computer program stored in the memory to implement the above-mentioned video interaction method based on user requirements.

[0049] For the video interaction method, device and equipment provided by the present application, first, obtain the target interaction request information sent by the target user through the corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user; secondly, based on the target location information, determine the target real-time monitoring video; then, transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device. Based on the above content, since after analyzing the target location information of the area of interest of the target user, the area of interest can be monitored in real time through the video monitoring device deployed in the area of interest reflected by the target location information, the obtained target real-time monitoring video has high real-time performance, thus ensuring that the obtained relevant information has high reliability. Therefore, the problem of relatively low reliability of video interaction existing in the prior art can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to make the above objects, features and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows.

[0051] Figure 1 It is a structural block diagram of the electronic device provided by the embodiment of the present application.

[0052] Figure 2 It is a schematic flowchart of the video interaction method based on user requirements provided by the embodiment of the present application.

[0053] Figure 3 It is a schematic diagram of the first updated user relationship network provided by the embodiment of the present application.

[0054] Figure 4 It is a schematic diagram of the local user relationship network provided by the embodiment of the present application.

[0055] Figure 5 It is a schematic block diagram of the video interaction device based on user requirements provided by the embodiment of the present application. Detailed implementation manners

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0057] Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0058] As Figure 1 shown, an embodiment of the present application provides an electronic device. Among them, the electronic device may include a memory, a processor, and a video interaction device based on user requirements.

[0059] Specifically, the memory and the processor are directly or indirectly electrically connected to achieve data transmission or interaction. For example, the memory and the processor may be electrically connected through one or more communication buses or signal lines. The video interaction device based on user requirements includes at least one software function module stored in the memory in the form of software or firmware. The processor is configured to execute the executable computer programs stored in the memory, such as the software function modules and computer programs included in the video interaction device based on user requirements, to implement the video interaction method based on user requirements provided by the embodiments of the present application.

[0060] Optionally, the memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0061] Moreover, the processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a System on Chip (SoC), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0062] It can be understood that Figure 1 The structure shown is only schematic, and the electronic device may also include more or fewer components than those shown Figure 1 herein, or have a configuration different from that shown Figure 1 For example, it may also include a communication unit for information interaction with other devices (such as drone devices, etc.).

[0063] Combined with Figure 2 , an embodiment of the present application further provides a video interaction method based on user needs applicable to the above-mentioned electronic device. Among them, the method steps defined by the processes related to the video interaction method based on user needs can be implemented by the electronic device.

[0064] Next, the Figure 2 specific process shown will be elaborated in detail.

[0065] Step S110: Obtain the target interaction request information sent by the target user through the corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user.

[0066] In an embodiment of the present application, the electronic device may obtain the target interaction request information sent by the target user through the corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user. Among them, the target location information is used to reflect the area of interest of the target user, such as a scenic area searched by the target user, etc.

[0067] Step S120: Determine a target real-time monitoring video based on the target location information.

[0068] In an embodiment of the present application, after obtaining the target location information, the electronic device may determine a target real-time monitoring video based on the target location information. The target real-time monitoring video is obtained by a video monitoring device deployed in the region of interest reflected by the target location information performing real-time video monitoring on the region of interest. That is to say, the electronic device may control the corresponding video monitoring device to perform video monitoring to obtain the corresponding target real-time monitoring video, and then, the target real-time monitoring video may be stored in the corresponding database.

[0069] Step S130: Transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device.

[0070] In an embodiment of the present application, after determining the target real-time monitoring video, the electronic device may transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device. Specifically, the target real-time monitoring video may be stored in the electronic device, and thus, the target real-time monitoring video may be directly transmitted to the target terminal device. Alternatively, the target real-time monitoring video may also be stored in other devices, and thus, the electronic device may send the storage address of the target real-time monitoring video to the target terminal device, so that the target terminal device may access and obtain the target real-time monitoring video based on the storage address.

[0071] Based on the above content, since after analyzing the target location information of the region of interest of the target user, the region of interest may be subjected to real-time video monitoring by a video monitoring device deployed in the region of interest reflected by the target location information, the obtained target real-time monitoring video has high real-time performance, thereby ensuring that the obtained relevant information has high reliability. Therefore, the problem of relatively low reliability of video interaction existing in the prior art can be improved.

[0072] It should be noted that for step S120, the specific manner of determining the target real-time monitoring video is not limited and may be selected according to actual requirements.

[0073] For example, in a first alternative embodiment, the video monitoring device may be a fixedly deployed device such as a camera to perform video monitoring on a fixed area. For another example, in a second alternative embodiment, the video monitoring device may be an airborne device of a drone to move within a region and perform corresponding video monitoring.

[0074] In a first alternative embodiment, in order to avoid problems such as resource waste that are prone to occur in video surveillance to a certain extent, step S120 described above may further include step S121 and step S122, and the specific content of each step is described as follows.

[0075] Step S121, identify each video surveillance device deployed in the region of interest reflected by the target location information to form at least one target video surveillance device.

[0076] In the embodiment of the present application, each video surveillance device deployed in the region of interest reflected by the target location information can be identified to form at least one target video surveillance device. That is to say, since the video surveillance device is fixedly deployed, therefore, each video surveillance device deployed in the region of interest reflected by the target location information (such as all the video surveillance devices of a certain scenic spot in a certain scenic area) can be directly used as the target video surveillance device.

[0077] Step S122, control each of the target video surveillance devices to start and conduct video surveillance respectively to obtain the target real-time surveillance video corresponding to each of the target video surveillance devices.

[0078] In the embodiment of the present application, after the target video surveillance device is identified, each of the target video surveillance devices can be controlled to start and conduct video surveillance respectively to obtain the target real-time surveillance video corresponding to each of the target video surveillance devices. That is to say, when there is no user demand, the video surveillance device can be not turned on, and only when there is a corresponding user demand, the video surveillance device is turned on to avoid resource consumption of the video surveillance device when there is no user demand.

[0079] It can be understood that in step S121 described above, the specific manner of identifying each video surveillance device deployed in the region of interest reflected by the target location information is not limited and can be selected according to actual needs. For example, in an alternative embodiment, in order to further reduce the degree of consumption of device resources, step S121 described above may further include step S121a, step S121b, and step S121c, and the specific content of each step is described as follows.

[0080] Step S121a, identify each video surveillance device deployed in the region of interest reflected by the target location information to form at least one candidate video surveillance device.

[0081] In the embodiment of the present application, each video surveillance device deployed in the region of interest reflected by the target location information can be determined to form at least one candidate video surveillance device. That is to say, each video surveillance device in the corresponding region can be used as a candidate video surveillance device.

[0082] Step S121b, when the number of devices of the at least one candidate video surveillance device is less than or equal to a pre-determined reference number, the at least one candidate video surveillance device is used as the target event surveillance device.

[0083] In the embodiment of the present application, after determining at least one candidate video surveillance device, when the number of devices of the at least one candidate video surveillance device is less than or equal to a pre-determined reference number, a corresponding candidate video surveillance device can be used as the target event surveillance device. That is to say, when the number of candidate video surveillance devices is small, there is no problem of excessive consumption of device resources. Therefore, each candidate video surveillance device can be directly used as the target event surveillance device. In addition, the specific value of the reference number is not limited, such as values of 1, 2, 3, 4, etc.

[0084] Step S121c, when the number of devices of the at least one candidate video surveillance device is greater than the reference number, based on a pre-determined device screening rule, at least one candidate video surveillance device is determined from the at least one candidate video surveillance device as the target event surveillance device.

[0085] In the embodiment of the present application, when the number of devices of the at least one candidate video surveillance device is greater than the reference number, based on a pre-determined device screening rule, at least one candidate video surveillance device is determined from the at least one candidate video surveillance device as the target event surveillance device. That is to say, when the number of candidate video surveillance devices is large, the at least one candidate video surveillance device can be further screened, so as to reduce the number of turned-on video surveillance devices to a certain extent to improve the problem of excessive consumption of device resources.

[0086] It can be understood that in the above step S121c, the specific manner of determining at least one candidate video surveillance device as the target event surveillance device from the at least one candidate video surveillance device is not limited and can be selected according to actual needs. For example, in an alternative embodiment, in order to improve the problem of excessive consumption of device resources and also take into account the adaptation to the actual needs of users to ensure the stickiness of users to the corresponding platform, the above step S121c can further include the following specific implementation content:

[0087] First, when the number of the at least one candidate video surveillance device is greater than the reference number, the area description information of the local surveillance area corresponding to each candidate video surveillance device can be obtained, where the area description information includes videos and / or texts, such as videos and texts of scenic spot introductions, etc., and the videos can be obtained by monitoring the scenic spots;

[0088] Second, the target user demand information sent by the target user through the target terminal device can be obtained, where the target user demand information is used to reflect the content of interest of the target user, such as the type of scenic spots the target user likes, etc., such as lakes, buildings, etc.;

[0089] Then, for each candidate video surveillance device, the matching degree between the area description information of the local surveillance area corresponding to the candidate video surveillance device and the target user demand information can be determined; Exemplarily, through a trained neural network, the area description information and the target user demand information can be respectively mined for features to obtain the corresponding area description information features and user demand information features, and then, the feature similarity (such as cosine similarity, etc.) between the area description information features and the user demand information features can be calculated as the corresponding matching degree. Specifically, for video frames, convolutional processing and self-attention processing can be performed, and for texts, word embedding processing and self-attention processing can be performed. The specific implementation process can refer to relevant existing technologies;

[0090] Finally, each candidate video surveillance device with a matching degree greater than or equal to a pre-determined reference matching degree can be used as a target event surveillance device, or the target number of candidate video surveillance devices with the largest matching degree can be used as target event surveillance devices; The specific values of the reference matching degree and the target number are not limited and can be configured according to actual needs.

[0091] In a second alternative embodiment, in order to ensure that the obtained target real-time surveillance video can better represent the corresponding area of interest, step S120 above can further include step S123 and step S124, and the specific content of each step is as described below.

[0092] Step S123, determining the video surveillance devices deployed in the area of interest reflected by the target location information to form target video surveillance devices.

[0093] In the embodiment of the present application, a video monitoring device deployed in the region of interest reflected by the target location information can be determined to form a target video monitoring device. Among them, the video monitoring device belongs to the airborne equipment of the unmanned aerial vehicle. That is to say, generally speaking, due to the characteristic that the airborne equipment of the unmanned aerial vehicle can move within a certain area, therefore, for the deployment of the airborne equipment of the unmanned aerial vehicle, the quantity requirement is not high. For example, one airborne equipment of the unmanned aerial vehicle can be deployed at one scenic spot, or multiple airborne equipment of the unmanned aerial vehicle can be deployed in one scenic area to accommodate the concurrent video interaction requirements. Based on this, the video monitoring device deployed in the region of interest reflected by the target location information can be directly used as the target video monitoring device.

[0094] Step S124, control the target video monitoring device to move in the region of interest reflected by the target location information, and perform video monitoring during the movement to obtain the corresponding target real-time monitoring video.

[0095] In the embodiment of the present application, after determining the target video monitoring device, the target video monitoring device can be controlled to move in the region of interest reflected by the target location information, and video monitoring is performed during the movement to obtain the corresponding target real-time monitoring video. That is to say, the target video monitoring device can be used to perform video monitoring on multiple local regions in the corresponding area, so that the obtained target real-time monitoring video has rich video content.

[0096] It can be understood that in the above step S124, the specific manner of controlling the target video monitoring device to move in the region of interest reflected by the target location information is not limited, and corresponding selection can be made according to actual needs. For example, in an alternative embodiment, in order to make the obtained target real-time monitoring video not only have rich video content but also be more adapted to the actual needs of the user, the above step S124 can further include step S124a, step S124b, and step S124c, and the specific content of each step is as follows.

[0097] Step S124a, obtain the target user demand information sent by the target user through the target terminal device.

[0098] In the embodiment of the present application, the target user demand information sent by the target user through the target terminal device can be obtained. Among them, the target user demand information is used to reflect the content of interest of the target user, such as favorite food, scenic spots, etc.

[0099] Step S124b, based on the target user demand information, determine a target movement path in the region of interest reflected by the target location information.

[0100] In an embodiment of the present application, after obtaining the target user requirement information, a target movement path may be determined in the area of interest reflected by the target location information based on the target user requirement information. Wherein, the target movement path includes at least a starting point and an ending point. That is to say, the areas in the target movement path are areas that the target user may be more interested in.

[0101] Step S124c, control the target video monitoring device to move along the target movement path in the area of interest reflected by the target location information, and perform video monitoring through the target video monitoring device during the movement to obtain a corresponding target real-time monitoring video.

[0102] In an embodiment of the present application, after determining the target movement path, the target video monitoring device may be controlled to move along the target movement path in the area of interest reflected by the target location information, and video monitoring may be performed through the target video monitoring device during the movement to obtain a corresponding target real-time monitoring video.

[0103] It can be understood that in the above step S124b, the specific manner of determining the target movement path in the area of interest reflected by the target location information is not limited and can be selected according to actual needs. For example, in an alternative embodiment, in order to ensure the reliability of determining the target movement path, the above step S124b may further include step b1, step b2, step b3, step b4, step b5, and step b6, and the specific content of each step is as follows.

[0104] Step b1, determine each historical user having requirement relevance with the target user based on the target user requirement information to form at least one relevant historical user.

[0105] In an embodiment of the present application, each historical user having requirement relevance with the target user may be determined based on the target user requirement information to form at least one relevant historical user (that is, the corresponding user requirement information has relevance with the target user requirement information). Wherein, the historical user refers to a user who has performed real-time video interaction with the electronic device in history.

[0106] Step b2, obtain the historical real-time monitoring videos formed by each of the relevant historical users in history to form at least one historical real-time monitoring video.

[0107] In an embodiment of the present application, after determining the relevant historical users, the historical real-time monitoring videos formed by each of the relevant historical users in history (that is, the corresponding target real-time monitoring videos in history) may be obtained to form at least one historical real-time monitoring video.

[0108] Step b3: Mine and aggregate multimodal features from the at least one historical real-time monitoring video and the target user requirement information to form multimodal user requirement features corresponding to the target user.

[0109] In an embodiment of the present application, after obtaining the historical real-time monitoring video, multimodal features can be mined and aggregated from the at least one historical real-time monitoring video and the target user requirement information to form multimodal user requirement features corresponding to the target user. Among them, the multimodal user requirement features are used to reflect the semantic information of the at least one historical real-time monitoring video and the target user requirement information. For example, for video frames, convolution processing can be performed to obtain corresponding convolution features, and then, for text, word embedding processing can be performed to obtain corresponding embedding features. In this way, the convolution features and the embedding features can be concatenated, or cross-attention processing can be performed on the embedding features based on the convolution features to obtain corresponding multimodal user requirement features.

[0110] Step b4: Perform path planning processing on the region of interest reflected by the target location information to form multiple candidate movement paths, and for each of the candidate movement paths, respectively obtain the node description information of each path node in the candidate movement path, and combine the node description information of each path node to form path description information corresponding to the candidate movement path.

[0111] In an embodiment of the present application, path planning processing can be performed on the region of interest reflected by the target location information to form multiple candidate movement paths, and for each of the candidate movement paths, respectively obtain the node description information (such as scenic spot introduction, etc.) of each path node in the candidate movement path, and combine the node description information of each path node to form path description information corresponding to the candidate movement path. Among them, the node description information includes video and / or text, that is, video introduction or text introduction, etc.

[0112] Step b5: Respectively perform feature mining on the path description information corresponding to each of the candidate movement paths to form path description features corresponding to each of the candidate movement paths.

[0113] In the embodiments of the present application, after obtaining the path description information, feature mining can be performed on the path description information corresponding to each of the candidate motion paths to form path description features corresponding to each of the candidate motion paths. Specifically, for a video frame, convolution processing can be performed to obtain corresponding convolution features. Then, for the text, word embedding processing can be performed to obtain corresponding embedding features. In this way, the convolution features and the embedding features can be concatenated, or cross-attention processing can be performed on the embedding features based on the convolution features to obtain path description features.

[0114] Step b6: Determine the similarity between the path description features corresponding to each of the candidate motion paths and the multi-modal user demand features respectively, and determine the candidate motion path with the maximum similarity as the target motion path.

[0115] In the embodiments of the present application, after obtaining the path description features and the multi-modal user demand features, the similarity (such as the cosine similarity between the features) between the path description features corresponding to each of the candidate motion paths and the multi-modal user demand features can be determined respectively, and the candidate motion path with the maximum similarity is determined as the target motion path. In addition, in the embodiments of the present application, the specific manifestation forms of the features are not limited. For example, they can be vectors. Specifically, for the text "Tianzi Mountain is famous for natural landscapes such as strange peaks, sea of clouds, sunrise, and waterfalls", through word segmentation and embedding processing, the following word embedding features corresponding to each word can be obtained:

[0116] Tianzi Mountain: [0.23, -0.12, 0.85, 0.09, -0.67, 0.54, 0.33, -0.44, 0.71, 0.11, 0.88, -0.12, -0.43,......, 0.56];

[0117] With: [-0.15, 0.21, 0.76, -0.02, 0.34, -0.56, 0.29, 0.33, -0.23, 0.57, -0.45, 0.61, 0.05,......, -0.11];

[0118] Strange peaks: [0.41, -0.67, 0.53, 0.22, 0.15, 0.78, -0.34, -0.59, 0.21, -0.44, 0.39, -0.25, 0.12,......, 0.67];

[0119] Sea of clouds: [0.37, 0.54, 0.65, -0.21, 0.56, -0.14, 0.68, 0.03, 0.44, -0.31, 0.72, 0.19, -0.22,......, -0.58];

[0120] Sunrise: [0.29, 0.11, 0.78, 0.23, -0.11, 0.91, -0.02, 0.57, 0.34, 0.65, 0.41, -0.55, 0.67,......, -0.23];

[0121] Waterfall: [0.62, 0.33, -0.09, 0.56, 0.79, 0.21, 0.54, -0.29, 0.63, 0.12, -0.67, 0.41, -0.13,......, 0.76];

[0122] Etc.: [-0.32, 0.48, 0.21, 0.67, -0.09, -0.17, 0.54, 0.77, -0.14, 0.33, 0.12, 0.43, 0.58,......, 0.29];

[0123] Nature: [0.68, 0.53, -0.23, 0.34, 0.74, 0.11, 0.44, -0.32, -0.07, 0.25, 0.33, 0.41, 0.58,......, 0.09];

[0124] Landscape: [0.45, 0.39, 0.57, 0.23, -0.21, 0.12, 0.88, 0.32, 0.47, -0.31, 0.51, 0.67, 0.19,......, -0.64];

[0125] Famous for: [0.22, 0.17, -0.12, 0.56, 0.19, 0.82, -0.33, 0.27, 0.75, 0.39, -0.23, 0.51, 0.41,......, 0.03].

[0126] It can be understood that in the above step b1, the specific manner of determining each historical user having demand relevance with the target user based on the target user demand information is not limited and can be selected according to actual needs. For example, in an alternative implementation manner, in order to ensure that the determined relevant historical users have high reliability and enable the demand of the target user to be characterized based on the historical real-time monitoring videos of the relevant historical users, the above step b1 may further include the following specific implementation content:

[0127] First, a historical user relationship network constructed based on each historical user and the historical user demand information corresponding to each historical user can be determined (in this historical user relationship network, there is demand relevance between two connected historical users), and, based on the target user and the target user demand information, the historical user relationship network is first updated to form a first-updated user relationship network. Among them, in the process of the first update, each historical user corresponding to the historical user demand information with a relevance (which can be the semantic similarity between texts, etc.) greater than or equal to a preset relevance (which can be configured according to the actual situation, such as 0.5, 0.6, 0.7, etc.) to the target user demand information is connected to the target user to represent that this historical user has demand relevance to the target user. Combining Figure 3 as shown;

[0128] Second, feature mining (such as word embedding processing and self-attention processing) can be performed on each of the to-be-processed user demand information, and the to-be-processed user demand features corresponding to each of the to-be-processed user demand information are output, where each of the to-be-processed user demand information belongs to the historical user demand information or the target user demand information;

[0129] Then, for each to-be-processed user in the first-updated user relationship network, in the first-updated user relationship network, a local user relationship network centered on this to-be-processed user is determined, where each of the to-be-processed users belongs to the historical user or the target user; Exemplarily, the connection degree between other to-be-processed users and the central to-be-processed user in the local user relationship network needs to be less than or equal to a preset value, such as numerical values 0, 1, 2, 3, etc. The connection degree can refer to the number of to-be-processed users included in the shortest connection path between two to-be-processed users. Taking Figure 3 as an example, when the target user is the center, other to-be-processed users with a connection degree of 0 include historical user 3, historical user 5, and historical user 8, other to-be-processed users with a connection degree of 1 include historical user 2 and historical user 7, other to-be-processed users with a connection degree of 2 include historical user 4, and other to-be-processed users with a connection degree of 3 include historical user 1 and historical user 6. Thus, if the preset value is 1, the local user relationship network centered on the target user is as Figure 4 shown;

[0130] After that, random walks can be performed in the first updated user relationship network to form multiple random walk paths (exemplarily, each random walk path is determined, as long as any two random walk paths are not exactly the same, or a random number of random walk paths can be randomly determined), and the multiple random walk paths are randomly serialized (i.e., the multiple random walk paths are arbitrarily sorted) to form a corresponding target path sequence, where the path end point of each random walk path is the target user, such as Figure 3 "Historical User 1, Historical User 4, Historical User 2, Historical User 3, Target User" in

[0131] Furthermore, according to the order of the multiple random walk paths in the target path sequence and the direction from the path start point to the path end point in the random walk path, the currently traversed user to be processed is determined, and based on the current user demand characteristics corresponding to each user to be processed in the local user relationship network of the currently traversed user to be processed, the current user demand characteristics corresponding to the currently traversed user to be processed are associated and mined to obtain corresponding associated user demand characteristics, and the associated user demand characteristics are used as the new current user demand characteristics to be processed. After the current user demand characteristics corresponding to the target user are associated and mined in the last random walk path, based on the feature similarity between the current user demand characteristics corresponding to each user to be processed, the first updated user relationship network is second updated to form a second updated user relationship network; exemplarily, two users to be processed with a feature similarity greater than or equal to the target similarity can be connected together, and the specific value of the target similarity can be configured according to actual needs, such as 0.5, 0.6, 0.7, etc.; for the associated mining, Figure 4For example, in the first step, based on the current user demand characteristics of the historical user 2 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the historical user 3 corresponding to the current user to be processed to obtain the new current user demand characteristics of the historical user 3. Also, based on the current user demand characteristics of the historical user 7 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the historical user 5 corresponding to the current user to be processed to obtain the new current user demand characteristics of the historical user 5. And, based on the current user demand characteristics of the historical user 7 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the historical user 8 corresponding to the current user to be processed to obtain the new current user demand characteristics of the historical user 8. In the second step, based on the current user demand characteristics of the historical user 3 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the target user corresponding to the current user to be processed to obtain the new current user demand characteristics of the target user. Also, based on the current user demand characteristics of the historical user 5 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the target user corresponding to the current user to be processed to obtain the new current user demand characteristics of the target user. And, based on the current user demand characteristics of the historical user 8 corresponding to the current user to be processed, cross-attention processing can be performed on the current user demand characteristics of the target user corresponding to the current user to be processed to obtain the new current user demand characteristics of the target user. In the third step, the new current user demand characteristics obtained from the three cross-attention processes can be averaged to obtain the actual current user demand characteristics of the target user.

[0132] Finally, in the second updated user relationship network, each historical user having a connection relationship with the target user can be determined to form at least one relevant historical user; that is, each historical user having a connection relationship with the target user in the second updated user relationship network can be used as the relevant historical user of the target user.

[0133] Combined with Figure 5 , the embodiment of the present application further provides a video interaction device based on user demand that can be applied to the above electronic device. Among them, the video interaction device based on user demand may include a location information determination module, a monitoring video determination module, and a monitoring video transmission module.

[0134] Specifically, the location information determination module can be used to obtain the target interaction request information sent by the target user through the corresponding target terminal device and parse the target interaction request information to obtain the target location information corresponding to the target user, where the target location information is used to reflect the area of interest of the target user. In the embodiment of the present application, the location information determination module can be used to executeFigure 2 For the step S110 shown, for the relevant content of the location information determination module, reference may be made to the description of step S110 above.

[0135] Specifically, the monitoring video determination module can be used to determine a target real-time monitoring video based on the target location information, where the target real-time monitoring video is obtained by a video monitoring device deployed in the region of interest reflected by the target location information for real-time video monitoring of the region of interest. In the embodiments of the present application, the monitoring video determination module can be used to execute Figure 2 For the step S120 shown, for the relevant content of the monitoring video determination module, reference may be made to the description of step S120 above.

[0136] Specifically, the monitoring video transmission module can be used to transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete real-time video interaction between the electronic device and the target terminal device. In the embodiments of the present application, the monitoring video transmission module can be used to execute Figure 2 For the step S130 shown, for the relevant content of the monitoring video transmission module, reference may be made to the description of step S130 above.

[0137] In the embodiments of the present application, corresponding to the above video interaction method based on user requirements applied to the electronic device, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium, and when the computer program runs, it executes each step of the video interaction method based on user requirements.

[0138] Among them, the steps executed when the foregoing computer program runs will not be elaborated one by one here, and reference may be made to the foregoing explanation of the video interaction method based on user requirements.

[0139] In summary, for the video interaction method, apparatus, and device based on user requirements provided in this application, first, target interaction request information sent by a target user through a corresponding target terminal device is obtained, and the target interaction request information is parsed to obtain target location information corresponding to the target user. Secondly, based on the target location information, a target real-time monitoring video is determined. Then, the target real-time monitoring video is transmitted to the target terminal device corresponding to the target user to complete real-time video interaction between the electronic device and the target terminal device. Based on the above content, since after analyzing the target location information of the area of interest of the target user, the area of interest can be monitored in real time through a video monitoring device deployed in the area of interest reflected by the target location information, the obtained target real-time monitoring video has high real-time performance, thereby ensuring that the relevant information obtained has high reliability. Therefore, the problem of relatively low reliability of video interaction in the existing technology can be improved.

[0140] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0141] In addition, each functional module in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0142] When the above-described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0143] The foregoing are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A video interaction method based on user requirements, characterized in that, Applied to an electronic device, wherein the video interaction method based on user needs includes: Obtain the target interaction request information sent by the target user through the corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user, wherein the target location information is used to reflect the area of interest of the target user; Determine the video surveillance device deployed in the area of interest reflected by the target location information to form a target video surveillance device, wherein the video surveillance device belongs to the airborne device of the unmanned aerial vehicle; obtain the target user demand information sent by the target user through the target terminal device, wherein the target user demand information is used to reflect the content of interest of the target user; based on the target user demand information, determine each historical user having demand relevance with the target user to form at least one relevant historical user, wherein the historical user refers to a user who has conducted real-time video interaction with the electronic device in history; obtain the historical real-time surveillance videos formed by each of the relevant historical users in history to form at least one historical real-time surveillance video; mine and aggregate the multi-modal features of the at least one historical real-time surveillance video and the target user demand information to form the multi-modal user demand features corresponding to the target user, wherein the multi-modal user demand features are used to reflect the semantic information of the at least one historical real-time surveillance video and the target user demand information; perform path planning processing on the area of interest reflected by the target location information to form a plurality of candidate movement paths, and for each of the candidate movement paths, respectively obtain the node description information of each path node in the candidate movement path, and combine the node description information of each path node to form the path description information corresponding to the candidate movement path, wherein the node description information includes video and / or text; respectively perform feature mining on the path description information corresponding to each of the candidate movement paths to form the path description features corresponding to each of the candidate movement paths; respectively determine the similarity between the path description features corresponding to each of the candidate movement paths and the multi-modal user demand features, and determine the candidate movement path with the maximum similarity as the target movement path; control the target video surveillance device to move along the target movement path in the area of interest reflected by the target location information, and perform video surveillance through the target video surveillance device during the movement to obtain the corresponding target real-time surveillance video; Transmit the target real-time surveillance video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device.

2. The video interaction method based on user requirements according to claim 1, wherein The step of determining each historical user having demand relevance with the target user based on the target user demand information to form at least one relevant historical user includes: Determine the historical user relationship network constructed based on each historical user and the historical user demand information corresponding to each historical user, and, based on the target user and the target user demand information, perform a first update on the historical user relationship network to form a first updated user relationship network. During the first update process, connect each historical user corresponding to the historical user demand information with a correlation degree greater than or equal to the preset correlation degree with the target user demand information to the target user, to represent that this historical user has demand relevance with the target user; Perform feature mining on each to-be-processed user demand information respectively, and output the to-be-processed user demand features corresponding to each to-be-processed user demand information, where each to-be-processed user demand information belongs to the historical user demand information or the target user demand information; For each to-be-processed user in the first updated user relationship network, determine a local user relationship network centered on this to-be-processed user in the first updated user relationship network, where each to-be-processed user belongs to the historical user or the target user; Perform random walks in the first updated user relationship network to form a plurality of random walk paths, and perform random serialization on the plurality of random walk paths to form a corresponding target path sequence, where the path end point of each random walk path is the target user; According to the order of the plurality of random walk paths in the target path sequence, and according to the direction from the path start point to the path end point in the random walk path, determine the currently traversed to-be-processed user, and, based on the currently to-be-processed user demand features corresponding to the to-be-processed users in the local user relationship network of the currently traversed to-be-processed user, perform association mining on the currently to-be-processed user demand features corresponding to the currently traversed to-be-processed user to obtain the corresponding associated user demand features, and use this associated user demand feature as the new currently to-be-processed user demand feature, and after performing association mining on the currently to-be-processed user demand features corresponding to the target user in the last random walk path, based on the feature similarity between the currently to-be-processed user demand features corresponding to each to-be-processed user, perform a second update on the first updated user relationship network to form a second updated user relationship network; In the second updated user relationship network, determine each historical user having a connection relationship with the target user to form at least one relevant historical user.

3. A video interaction device based on user requirements, characterized in that, Applied to an electronic device, where the video interaction device based on user demand includes: A location information determination module, configured to obtain a target interaction request information sent by a target user through a corresponding target terminal device, and parse the target interaction request information to obtain the target location information corresponding to the target user, where the target location information is used to reflect the area of interest of the target user; A monitoring video determination module, configured to determine a video monitoring device deployed in the region of interest reflected by the target location information to form a target video monitoring device, where the video monitoring device belongs to an airborne device of a drone; obtain target user requirement information sent by the target user through the target terminal device, where the target user requirement information is used to reflect the content of interest of the target user; determine each historical user having a requirement relevance with the target user based on the target user requirement information to form at least one relevant historical user, where the historical user refers to a user who has had a real-time video interaction with the electronic device in history; obtain historical real-time monitoring videos formed by each of the relevant historical users in history to form at least one historical real-time monitoring video; perform mining and aggregation of multi-modal features on the at least one historical real-time monitoring video and the target user requirement information to form multi-modal user requirement features corresponding to the target user, where the multi-modal user requirement features are used to reflect the semantic information of the at least one historical real-time monitoring video and the target user requirement information; perform path planning processing on the region of interest reflected by the target location information to form a plurality of candidate movement paths, and for each of the candidate movement paths, respectively obtain the node description information of each path node in the candidate movement path, and combine the node description information of each path node to form path description information corresponding to the candidate movement path, where the node description information includes video and / or text; respectively perform feature mining on the path description information corresponding to each of the candidate movement paths to form path description features corresponding to each of the candidate movement paths; respectively determine the similarity between the path description features corresponding to each of the candidate movement paths and the multi-modal user requirement features, and determine the candidate movement path with the maximum similarity as the target movement path; control the target video monitoring device to move along the target movement path in the region of interest reflected by the target location information, and perform video monitoring through the target video monitoring device during the movement to obtain a corresponding target real-time monitoring video; A monitoring video transmission module, configured to transmit the target real-time monitoring video to the target terminal device corresponding to the target user to complete the real-time video interaction between the electronic device and the target terminal device.

4. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor connected to the memory, configured to execute the computer program stored in the memory to implement the video interaction method based on user requirements according to any one of claims 1-2.

Citation Information

Patent Citations

  • Method and system for processing content of aerial photographic service of unmanned aerial vehicle

    CN105824967A

  • Live broadcast system capable of acquiring high-definition pictures immediately

    CN108696724A