Face recognition method, computer device, medium and product
By obtaining and analyzing user equipment, browser and environment information, determining the best camera configuration and backup strategy, the compatibility and accuracy of face recognition on different platforms and devices are solved, and efficient and stable face recognition effects are achieved.
Patent Information
- Application Number
- CN202510320981.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
It is difficult for the prior art to achieve consistent, efficient and stable facial recognition on different platforms, devices and browsers.
By obtaining user device information, browser information and environment status, we can determine whether the browser supports the camera function and determine whether there is an available camera based on the device information. If not, the backup strategy for uploading face images can be triggered; if not, the optimal camera configuration can be determined and face image acquisition can be performed.
Face recognition that works normally in various scenarios is realized, the scope of application is expanded, and the accuracy and performance of recognition is improved by optimizing the camera configuration.
Smart Images

Figure CN120148091A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of face recognition technology, and in particular, to a face recognition method, a computer device, a medium, and a product. Background Art
[0002] With the wide application of face recognition technology, especially in scenarios such as identity verification and payment, the face recognition solution on mobile devices has gradually become a key function. Currently, there are mainly two implementation methods: H5 Web applications and WeChat mini-programs. H5 applications call the camera through the browser and process images. However, different browsers have different degrees of support for Web APIs (such as camera calls), resulting in the inability to be compatible with all browsers. In addition, the cameras and processing capabilities of different devices vary greatly, affecting the accuracy and performance of recognition. WeChat mini-programs provide more encapsulated API interfaces, which simplifies development. However, the development of WeChat mini-programs must follow specific specifications and restrictions of WeChat, with relatively low flexibility and freedom, and can only run on the WeChat platform, without cross-platform universality. Therefore, how to efficiently and stably implement face recognition under different platforms is an important issue in development. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a face recognition method, a computer device, a medium, and a product, which can solve the problem in the prior art that due to device differences and browser compatibility issues, it is impossible to achieve consistent, efficient, and stable face recognition on different platforms, different devices, and different browsers.
[0004] In a first aspect, embodiments of the present disclosure provide a face recognition method, adopting the following technical solution:
[0005] In response to a face recognition start instruction, obtain user device information, browser information, and environmental status;
[0006] According to the browser information, determine whether the current browser supports the camera function. If so, according to the user device information, determine whether there is an available camera. If there is no available camera, trigger a first backup strategy to upload a face picture and perform collection based on the uploaded face picture;
[0007] If there is an available camera, obtain the available camera, and according to the browser information, determine the range of camera parameters supported by the browser; the range of camera parameters includes the range of video stream width, the range of video stream height, the range of supported frame rates, and the range of camera directions;
[0008] According to the available camera, the range of camera parameters supported by the browser, and the environmental status, determine the optimal camera configuration;
[0009] Call the best camera configuration to obtain face image information, and process the face image information based on a preset image analysis strategy to obtain a face recognition result.
[0010] Optionally, if the current browser does not support the camera function, trigger a missing filling strategy, and perform a filling library loading process on the current browser based on the missing filling strategy, so that the processed current browser supports the camera function;
[0011] The triggering of the missing filling strategy and the performance of the filling library loading process on the current browser based on the missing filling strategy, so that the processed current browser supports the camera function, includes:
[0012] Dynamically create a script element, which is used to load a filling library that can fill the missing camera function of the current browser later;
[0013] Set the source address attribute of the script element, where the source address points to the network link of the filling library used to fill the missing camera function of the current browser;
[0014] Add a load completion event handling to the script element. After the script element is successfully loaded, the processed current browser supports the camera function.
[0015] Optionally, the triggering of the first backup strategy for face picture uploading includes:
[0016] Provide a file selector to allow the user to select a picture file and limit the file type to the picture format;
[0017] After the user selects a file, check whether the file size meets the preset requirements. If it exceeds the limit, prompt the user to select a smaller file;
[0018] After the user selects a picture, display a picture preview for the user to confirm;
[0019] After the user confirms, preprocess the picture to be uploaded, and the preprocessing includes one or more of compression and cropping;
[0020] Create a FormData object, and encapsulate the preprocessed picture and necessary form data into the FormData object; the necessary form data includes one or more of user ID, request identifier, and image processing options;
[0021] Set the request headers corresponding to the FormData object;
[0022] Construct a fetch request and initiate a POST request using the fetch method. The POST request is used to send the FormData object as the request body to the backend API.
[0023] Optionally, if the current browser does not support the camera function, trigger the first fallback strategy for face image recognition.
[0024] Optionally, the obtaining of the available camera includes:
[0025] Obtain all cameras and their unique identifiers;
[0026] Perform a permission test on each camera to obtain the available camera.
[0027] Optionally, the processing of the face image information based on a preset image analysis strategy to obtain a face recognition result includes:
[0028] Locate the target face area based on a preset detection strategy;
[0029] Extract the original feature points in each frame of the target face area, and fuse the original feature points of consecutive frames to obtain target feature points;
[0030] Process the target feature points through a neural network to generate a face feature vector;
[0031] According to a preset vector library, obtain the matching similarity of the face feature vector;
[0032] If the matching similarity is less than a preset threshold, output a face recognition success instruction;
[0033] If the matching similarity is not less than a preset threshold, trigger the second fallback strategy to re-collect, or trigger the first fallback strategy to upload a face image, and perform recognition based on the uploaded face image.
[0034] Optionally, the locating of the target face area based on a preset detection strategy includes:
[0035] According to the best camera configuration, perform a quick scan on a low-resolution frame. If no face is detected, switch to a high-resolution frame for accurate scanning to obtain the target face;
[0036] When multiple faces are detected, select the face closest to the center of the camera as the target face;
[0037] Based on the target face, determine the target boundary, and the target boundary does not contain redundant background information;
[0038] The target face and the target boundary constitute the target face area.
[0039] In a second aspect, the embodiments of the present disclosure also provide a computer device, which adopts the following technical solutions:
[0040] The computer device includes:
[0041] at least one processor; and,
[0042] a memory communicatively connected to the at least one processor; wherein,
[0043] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the face recognition method described in any one of the above.
[0044] In a third aspect, an embodiment of the present disclosure further provides a computer-readable storage medium that stores computer instructions for causing a computer to execute the face recognition method described in any one of the above.
[0045] In a fourth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0046] The face recognition method disclosed in this application, in response to a face recognition start instruction, obtains user device information, browser information, and environmental status; determines whether the current browser supports the camera function according to the browser information, and if so, determines whether there is an available camera according to the user device information. If there is no available camera, triggers a first backup strategy to upload a face picture and collect based on the uploaded face picture; if there is an available camera, obtains the available camera and determines the range of camera parameters supported by the browser according to the browser information; determines the best camera configuration according to the available camera, the range of camera parameters supported by the browser, and the environmental status; calls the best camera configuration to obtain face image information, and processes the face image information based on a preset image analysis strategy to obtain a face recognition result; this method takes into account the differences between different user devices and browsers, as well as the influence of the environmental status, provides a backup strategy by judging the browser support situation and the availability of the device camera, enables the face recognition method to work properly in multiple scenarios, and expands the applicable range; determines the best camera configuration according to the environmental status and the parameter range supported by the browser, and the quality of the collected face images is higher. Combined with the preset image analysis strategy, it can effectively improve the accuracy of face recognition; provides a backup method for face picture upload for users without an available camera, avoids the situation where the face recognition function cannot be used due to device hardware problems, and ensures that users can complete the recognition operation smoothly.
[0047] The above description is only an overview of the technical solution of the present disclosure. In order to better understand the technical means of the present disclosure, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings as follows. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0049] Figure 1 It is a schematic flowchart of the face recognition method provided by the embodiment of the present disclosure.
[0050] Figure 2 It is a schematic flowchart of the method for loading a fill library for the current browser based on a missing filling strategy provided by the embodiment of the present disclosure.
[0051] Figure 3 It is a schematic flowchart of the method for triggering the first backup strategy to upload a face picture provided by the embodiment of the present disclosure.
[0052] Figure 4 It is a schematic flowchart of the method for obtaining available cameras provided by the embodiment of the present disclosure.
[0053] Figure 5 It is a schematic flowchart of the method for processing face image information based on a preset image analysis strategy to obtain a face recognition result provided by the embodiment of the present disclosure.
[0054] Figure 6 It is a schematic flowchart of the method for locating the target face area provided by the embodiment of the present disclosure.
[0055] Figure 7 It is a schematic structural diagram of a computer device provided by the embodiment of the present disclosure. Detailed Embodiments
[0056] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0057] It should be clear that the following uses specific specific examples to illustrate the implementation manners of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope protected by the present disclosure.
[0058] It should also be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement a device and / or practice a method. Additionally, this device and / or this method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.
[0059] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically, and only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0060] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0061] Referring to Figure 1 , the first aspect of this application discloses a face recognition method, including:
[0062] S100, in response to a face recognition start instruction, obtain user device information, browser information, and environmental status.
[0063] For example, when the user clicks the "Start Face Recognition" button on the web page, the system receives this startup instruction. Through front-end technologies such as JavaScript, the system obtains the user's device information, such as device models (such as iPhone 15, Huawei P70, etc.), operating system versions (such as iOS 17, Windows 11, etc.); browser information, including browser types (such as Chrome, Firefox, etc.) and version numbers; the environmental status can be obtained through device sensor data such as light sensors and gyroscopes, for example, to determine whether the current environmental light is bright, dim, or at night, and whether the device is in landscape or portrait orientation.
[0064] In this step, the user's device information helps to subsequently determine whether the device has the hardware conditions for face recognition. The camera performance and configuration of different devices are different. Browser information can help determine whether the browser supports camera calls and the scope of supported functions. The environmental status information can provide a basis for subsequently adjusting the camera parameters to adapt to different environments and improve the accuracy of face recognition.
[0065] Furthermore, the user's device information includes the user's hardware performance; the browser information includes one or more of the browser type and browser version; the environmental status includes one or more of the network type, latency, and broadband.
[0066] S200, based on the browser information, determine whether the current browser supports the camera function. If so, based on the user's device information, determine whether there is an available camera. If there is no available camera, trigger the first backup strategy to upload a face picture and perform collection based on the uploaded face picture.
[0067] The API provided by the browser can be used to determine whether the camera function is supported. For example, in JavaScript, the navigator.mediaDevices.getUserMedia method can be used for testing. If this method is available, it means that the browser supports the camera function. Then, based on the device information, query the system settings or call relevant interfaces to determine whether there is an available camera on the device. For example, a laptop may have a built-in camera, and a mobile phone has a front and rear camera. If there is no available camera, the system will pop up a prompt box to guide the user to select a face picture from the local album for upload, and then perform preprocessing such as cropping and noise reduction on the uploaded picture to extract face features.
[0068] In this step, judging the browser support situation first can avoid invalid camera call operations on browsers that do not support the camera. For devices without an available camera, providing an alternative way to upload face pictures ensures the availability of the face recognition function and expands the scope of application.
[0069] S300, if there is an available camera, obtain the available camera and determine the range of camera parameters supported by the browser according to the browser information; the range of camera parameters includes the video stream width range, the video stream height range, the supported frame rate range, and the camera orientation range.
[0070] Obtain the list of available cameras through the browser's MediaDevices API. For example, in the Chrome browser, the navigator.mediaDevices.enumerateDevices method can be used. Then, determine the range of camera parameters it supports according to the information provided by the browser's documentation and API. For example, the video stream width range supported by some browsers is 640 - 1920 pixels, the height range is 480 - 1080 pixels, the frame rate range is 15 - 60 frames per second, and the camera orientation can be front or rear.
[0071] In this step, clarifying the range of camera parameters supported by the browser provides a basis for subsequent selection of the best camera configuration, ensuring that appropriate parameters can be selected in different browser environments and improving the quality of face image acquisition.
[0072] S400, determine the best camera configuration according to the available camera, the range of camera parameters supported by the browser, and the environmental status.
[0073] For example, if the ambient light is dim, preferentially select a camera configuration that supports high frame rates and high sensitivity to reduce image blur and noise. If the device is in landscape mode, select a configuration with a larger video stream width. Considering the performance of the available camera, the parameter range supported by the browser, and the environmental status comprehensively, calculate the optimal video stream width, height, frame rate, and camera orientation through a preset algorithm. For example, in a dimly lit night environment, select a front camera configuration with a frame rate of 30 frames per second, a video stream width of 1280 pixels, and a height of 720 pixels.
[0074] In this step, determining the best camera configuration according to the actual situation can capture the highest-quality face images under different devices, browsers, and environmental conditions, thereby improving the accuracy and reliability of face recognition.
[0075] S500, call the best camera configuration to obtain face image information, and process the face image information based on a preset image analysis strategy to obtain a face recognition result.
[0076] Obtain high-quality face images through the best camera configuration and process them through a preset image analysis strategy, which can accurately and quickly obtain face recognition results, improving the accuracy and efficiency of recognition.
[0077] The face recognition method disclosed in this application takes into account the differences between different user devices and browsers, as well as the impact of the environmental state. By judging the browser support situation and the availability of the device camera, it provides alternative strategies, enabling the face recognition method to work properly in various scenarios and expanding the scope of application; determines the optimal camera configuration according to the environmental state and the parameter range supported by the browser, and the quality of the captured face images is higher. Combined with the preset image analysis strategy, it can effectively improve the accuracy of face recognition; provides an alternative way for users without available cameras to upload face pictures, avoiding the situation where the face recognition function cannot be used due to device hardware problems, and ensuring that users can successfully complete the recognition operation.
[0078] In a specific embodiment, if the current browser does not support the camera function, a missing filling strategy is triggered, and the current browser is processed for filling library loading based on the missing filling strategy, and the processed current browser supports the camera function.
[0079] Refer to Figure 2 , the method for processing the current browser for filling library loading based on the missing filling strategy specifically includes:
[0080] S201, dynamically create a script element, which is used to load a filling library that can fill the missing camera function of the current browser subsequently.
[0081] Among them, dynamically creating a script element means generating a new script element in a programming manner in the browser page running environment under the condition that it is determined that the current browser does not support the camera function.
[0082] When it is detected through the program that the current browser does not support the camera function, in the environment where the browser page runs, a new script element is generated by using programming means. For example, it is like creating a special "tool" (script element) in a virtual "workshop" (browser page running environment) according to certain rules (programming method), and this "tool" will be used to load a filling library that can fill the missing camera function of the current browser.
[0083] Through this step, it has high flexibility. It is not necessary to fixedly write the code for loading the script in the web page file at the beginning, but to decide whether to create and load the script according to whether the browser actually supports the camera function; if the browser itself supports the camera function, there is no need to create and load this script, avoiding unnecessary resource waste; if the browser does not support it, this "tool" can be created in time to solve the problem.
[0084] S202, set the source address attribute of the script element, where the source address points to the network link of the filling library used to fill the missing camera function of the current browser.
[0085] Among them, the polyfill library pointed to by the source address is prepared specifically to address the deficiencies in the camera function of the current browser, and its network link is a specific link that can provide the corresponding filling function.
[0086] Specifically, after creating the script element, its source address property needs to be set for this script element. This source address is like a "navigation" that points to the network link where the polyfill library prepared specifically to address the deficiencies in the camera function of the current browser is located. For example, this link may be the address of a specific polyfill library file stored on a certain server.
[0087] By setting the source address property, the browser can accurately find the location of the polyfill library. Just like when you want to go to a specific place to get something, having the accurate address allows you to arrive smoothly. In this way, the browser can download and load the polyfill library from the specified network link, and use the professional and optimized polyfill library to address the deficiencies in the browser's function, making the function implementation more reliable and having better compatibility.
[0088] S203, add a load completion event handler to the script element. After the script element is successfully loaded, the current browser supports the camera function after processing.
[0089] Specifically, load the target polyfill from the preset Polyfill library, aiming to simulate the API of modern browsers so that old browsers can also support the camera function. The specific operation is as follows: First, create a script tag, just like creating a script element before. Then set the source address property of this script tag to the network link of the Polyfill library, such as a specific website address. Next, add a load completion event handler to this script tag. After the script tag is successfully loaded, try to access the camera again. If the access is successful, the captured images of the camera can be processed normally; if an error occurs during the loading process, such as a network problem preventing the download of the polyfill library, an error message indicating the loading failure will be output, which is convenient for developers to know where the problem lies.
[0090] Adding a load completion event handler can ensure that the camera function is only tried after the polyfill library is successfully loaded; by trying to access the camera again in the load completion event, it can be verified whether the polyfill library has taken effect, ensuring that the browser can truly support the camera function after loading the polyfill library, and the error handling mechanism can send a signal in a timely manner when the polyfill library fails to load, which is convenient for developers to debug and solve problems.
[0091] The method for loading a filling library for the current browser based on the missing filling strategy disclosed in S201 - S203 enables old browsers that originally do not support the camera function to support the camera by loading the filling library. This expands the applicable scope of applications that rely on the camera function, such as face recognition, allowing more users using old browsers to also use relevant services normally; the filling library exists in the form of an independent network link. When the filling library is updated, only the filling library file pointed to by this link needs to be updated, without modifying a large amount of web page code, which greatly reduces the cost and difficulty of maintenance; for users using browsers that do not support the camera function, they can also use functions that rely on the camera normally and will not be unable to use relevant services due to browser compatibility issues, thus enhancing the user experience during use and making users feel more convenient and smooth.
[0092] In another embodiment, if the current browser does not support the camera function, Polyfill or Shim technology can be used to support old - version browsers by supplementing missing APIs; or, a cross - browser compatibility testing framework (such as BrowserStack or Sauce Labs) can be adopted to automatically test the functional performance in different browser environments; or, the encapsulation ability of a front - end framework (such as React or Vue.js) can be introduced to reduce compatibility issues directly relying on native APIs.
[0093] Refer to Figure 3 , the method for triggering the first backup strategy for face picture upload specifically includes:
[0094] A110, Provide a file selector to allow users to select picture files and limit the file type to picture format.
[0095] The picture formats include JPEG and PNG.
[0096] For example, a button or a specific area can be set on the web page interface, and when the user clicks, a file selector will pop up; in the file selector, only picture files in JPEG and PNG formats are displayed for the user to select, and other types of files are not displayed.
[0097] Limiting the file type to picture format can ensure that the files selected by users meet the system requirements, avoid errors in subsequent processing caused by users selecting non - picture files, and improve the stability and processing efficiency of the system.
[0098] A120, After the user selects a file, check whether the file size meets the preset requirements. If it exceeds the limit, prompt the user to select a smaller file.
[0099] Among them, meeting the preset requirements can be not exceeding 5MB.
[0100] When the user selects a picture file, the system will automatically obtain the size information of the file and compare it with the preset size limit of 5MB; if the file size exceeds 5MB, the system will pop up a prompt box to inform the user that the file is too large and ask the user to select a smaller file.
[0101] Limiting the file size can avoid problems such as network congestion and slow server processing caused by uploading overly large files, ensuring the smoothness and efficiency of the upload process. At the same time, it also helps to control the use of server storage space.
[0102] A130, after the user selects a picture, display a picture preview for the user to confirm.
[0103] When the user selects a picture, the system will display the preview effect of the selected picture in a specific area on the web page, allowing the user to intuitively see the content of the picture they selected.
[0104] Providing the picture preview function allows the user to confirm whether the selected picture is correct, avoiding the upload of incorrect pictures due to misselection, which enhances the interaction between the user and the system and improves the user experience.
[0105] A140, after the user confirms, preprocess the picture to be uploaded. The preprocessing includes one or more of compression and cropping.
[0106] Specifically, the system can use the Canvas API or the third-party library Compressor.js to compress the picture. For example, when using the Canvas API, the system will load the picture into the Canvas and compress it by adjusting the quality parameters of the picture, etc., to reduce the file size; according to business requirements, the system will provide a cropping tool, and the user can use this tool to crop the picture, and then upload it after cropping.
[0107] Compressing the picture can significantly reduce the file size, thereby accelerating the upload speed and saving network bandwidth; the cropping function allows the user to adjust the picture according to actual needs to ensure that the uploaded picture meets the business requirements.
[0108] A150, create a FormData object and encapsulate the preprocessed picture and necessary form data into the FormData object.
[0109] Among them, the necessary form data includes one or more of the user ID, request identifier, and image processing options.
[0110] Specifically, create a FormData object and then add the pre - processed image to this object. At the same time, necessary form data such as user ID, request identifier, image processing options, etc. will also be added to the FormData object.
[0111] Using the FormData object can conveniently encapsulate the image and form data together, simplifying the data processing and transmission process. It can handle the mixed transmission of binary data (such as images) and text data (such as form information) well.
[0112] A160, set the request headers corresponding to the FormData object.
[0113] Specifically, set the corresponding request header information according to the characteristics of the FormData object to ensure that the server can correctly identify and process the request.
[0114] Correctly setting the request headers can ensure that the server can accurately parse the data in the request body, avoiding problems such as data transmission failure or the server being unable to correctly process the data due to improper request header settings.
[0115] For example, when using the `fetch` API, ensure that the request headers contain the correct `Content - Type` (usually `multipart / form - data`), and set other necessary header information (such as authentication Token) according to the requirements of the backend.
[0116] A170, construct a fetch request and use the fetch method to initiate a POST request. The POST request is used to send the FormData object as the request body to the backend API.
[0117] Specifically, construct a fetch request, specify the request method as POST, and use the previously encapsulated FormData object as the request body; then use the fetch method to send this request to the backend API.
[0118] Using the fetch method to initiate a POST request is a modern, flexible and powerful way to interact with the backend server. It can conveniently handle requests and responses, support asynchronous operations, and improve the performance and response speed of the system.
[0119] In this embodiment, ensure that the backend server supports Cross - Origin Resource Sharing (CORS). If the backend does not directly support CORS, a proxy server (such as Nginx) can be used to forward the request.
[0120] The method for triggering the first backup strategy to upload face pictures disclosed in A110 - A170. Through the entire solution, by providing functions such as file selection restrictions, picture preview, and cropping tools, users can more conveniently and accurately select and process the pictures to be uploaded, enhancing the interactivity between users and the system and improving the user experience. Compressing the pictures can reduce the file size, speed up the upload speed, save network bandwidth, and at the same time reduce the processing burden on the server. Reasonable file size limits and data encapsulation methods also contribute to improving the stability and processing efficiency of the system. By restricting file types, adding necessary form data, and correctly setting the request headers, it can ensure that the uploaded data is accurate and can be securely transmitted to the backend server, guaranteeing the normal operation of the business.
[0121] In this embodiment, before sending the actual request, the browser will send an `OPTIONS` request (preflight request) to confirm whether the server allows cross - origin requests. Ensure that the backend correctly processes the preflight request and returns an allowed response header.
[0122] The `Promise` chain of the `fetch` API can be combined with the `progress` event of `ReadableStream` or `XMLHttpRequest` to listen to the upload progress and provide real - time feedback to the user.
[0123] Furthermore, this application also includes: providing an automatic or manual retry mechanism that allows users to re - upload pictures after failure; displaying a prompt message such as "Uploading" or "Please wait" during the upload process to prevent users from making incorrect operations; after the upload is completed, prompting the user "Upload successful" and providing guidance for the next operation (such as returning to the list page or continuing to upload); also providing a button to cancel the upload, allowing users to cancel the operation at any time during the upload process.
[0124] In another embodiment, if the current browser does not support the camera function, trigger the first backup strategy for face picture recognition.
[0125] Refer to Figure 4 , the method for obtaining available cameras specifically includes:
[0126] S301, obtain all cameras and their unique identifiers.
[0127] In a specific embodiment, in modern web development, the navigator.mediaDevices.enumerateDevices() method is usually used to obtain information about all media devices on a device. After calling this method, a Promise object is returned. When this Promise is resolved, an array containing information about all media devices is obtained. By filtering this array and only retaining devices with a device type of "videoinput", information about all cameras can be obtained; each camera information object will contain a deviceId property, which is the unique identifier of the camera.
[0128] After obtaining all cameras and their unique identifiers, developers can effectively manage the cameras. For example, different cameras can be distinguished based on the unique identifier, which facilitates subsequent operations on specific cameras. At the same time, it provides a basis for subsequent function expansion. For example, in a multi-camera device, different cameras can be selected for use according to user needs, or multiple cameras can be controlled simultaneously, etc.
[0129] S302, perform a permission test on each camera to obtain available cameras.
[0130] Specifically, for each obtained camera, a preset method can be used to request access to the camera. For example, the navigator.mediaDevices.getUserMedia() method can be used to request access to the camera. When calling this method, a constraint object is passed in, which specifies the unique identifier of the camera to be used; if the request is successful, it means the camera is available; if the request fails, it means the camera is unavailable or the user has refused the access permission.
[0131] Through the permission test, it can be accurately known which cameras are truly available, avoiding errors caused by unavailable cameras during subsequent use; performing a permission test before using the camera can prompt the user in advance whether to allow access to the camera, giving the user a better sense of control, and at the same time avoiding abnormal situations after the user refuses the permission.
[0132] In this embodiment, this solution can accurately find available cameras by first obtaining information about all cameras and then performing a permission test on each camera, improving the reliability of the system in using cameras; it can adapt to different devices and browser environments, ensuring that available cameras can be correctly obtained in various situations, effectively avoiding unexpected errors when using cameras, handling permission issues in advance, and enabling users to use camera-related functions more smoothly.
[0133] Refer to Figure 5, for the method of "processing the face image information based on a preset image analysis strategy to obtain a face recognition result" in S500, it specifically includes:
[0134] A100, locate the target face area based on a preset detection strategy.
[0135] Specifically refer to Figure 6 , the method for locating the target face area specifically includes:
[0136] A110, according to the optimal camera configuration, perform a quick scan on the low-resolution frame. If no face is detected, switch to the high-resolution frame for precise scanning to obtain the target face.
[0137] First, according to the optimal camera configuration parameters (such as resolution, frame rate, etc.), obtain the video stream from the camera. For each frame in the video stream, convert it into a low-resolution version. For example, downsample a frame originally with a resolution of 1920x1080 to 320x240. Use a lightweight face detection algorithm (such as Haar cascade classifier-based) to perform a quick scan on the low-resolution frame. This algorithm can quickly process the image and determine whether there is a face. If no face is detected on the low-resolution frame, use the original high-resolution frame and switch to a more precise but computationally intensive face detection algorithm (such as the deep learning-based MTCNN algorithm) for scanning to ensure that the target face is not missed.
[0138] In this step, the scanning speed of the low-resolution frame is fast, which can complete the preliminary detection in a short time and reduce unnecessary computational overhead. Only when no face is detected on the low-resolution frame is the high-resolution frame used, avoiding the performance bottleneck caused by always using the high-resolution frame; the high-resolution frame combined with a precise detection algorithm can accurately identify the target face in the case of failure to detect on the low-resolution frame, improving the accuracy of face detection.
[0139] A120, when multiple faces are detected, select the face closest to the center of the camera as the target face.
[0140] After multiple faces are detected, obtain the center coordinates of each face (usually represented by the center point of the face detection box). Calculate the center coordinates of the camera image frame. For example, for an image frame of 1920x1080, the center coordinates are (960, 540). Calculate the distance between the center coordinates of each face and the center coordinates of the camera image frame (the Euclidean distance formula can be used). Select the face with the smallest distance as the target face.
[0141] In a multi-person scenario, the face closest to the center of the camera is usually the main object of concern. Selecting this face as the target face meets the general needs of users, enables more precise subsequent processing of key figures, can effectively avoid complex selection logic among multiple faces, and improves the processing efficiency and stability of the system.
[0142] A130. Based on the target face, determine the target boundary, and the target boundary does not contain redundant background information.
[0143] The target face and the target boundary constitute the target face region.
[0144] After determining the target face, first obtain the position and size information of the face detection box. The position of the facial features of the face can be determined using a key-point detection algorithm (such as the 68 key-point detection in the dlib library); based on the position of the facial features and the face detection box, fine-tune the detection box to remove redundant background information around it. For example, by expanding or shrinking the boundary of the detection box to just enclose the key parts of the face, such as the forehead, cheeks, chin, etc.
[0145] In this step, after removing the redundant background information, the target face region is more pure, which is beneficial to subsequent face feature extraction, recognition, etc., improving the accuracy and reliability of the processing; a smaller target boundary means a reduction in the amount of data to be processed, thereby reducing the computational cost and improving the processing speed of the system.
[0146] A200. Extract the original feature points in each frame of the target face region, and fuse the original feature points of consecutive frames to obtain the target feature points.
[0147] By fusing the feature points of consecutive frames, the influence caused by single-frame image noise or pose changes can be reduced, improving the stability of the feature points; integrating the information of multiple frames can more comprehensively reflect the features of the face, helping to improve the accuracy of subsequent recognition.
[0148] Specifically, the method for obtaining the target feature points includes:
[0149] A201. Compare the original feature points of consecutive frames, eliminate the jitter values in single-frame detection, and obtain the first feature points.
[0150] A202. Smooth the first feature points within consecutive frames to obtain the target feature points.
[0151] Furthermore, the glasses area and the mouth area can also be extracted based on the target feature points for local contrast enhancement.
[0152] A300. Process the target feature points through a neural network to generate a face feature vector.
[0153] Specifically, a pre-trained face recognition neural network such as FaceNet or ArcFace can be used. These networks are trained with a large amount of data and can map facial feature points to a low-dimensional feature vector space. The target feature points are input into the selected neural network, and through the forward propagation calculation of the network, a face feature vector is output.
[0154] In this embodiment, the high-dimensional feature point information is compressed into a low-dimensional feature vector, reducing the data volume while retaining the key features of the face. The neural network can learn the deep features of the face, and the generated feature vector has good discrimination, which is beneficial to subsequent similarity calculation and recognition.
[0155] A400. Obtain the matching similarity of the face feature vector according to the preset vector library.
[0156] Extract the feature vectors of the face images of known persons in advance and store them in a vector library. Use methods such as cosine similarity or Euclidean distance to calculate the similarity between the current face feature vector and each feature vector in the vector library.
[0157] By calculating the similarity, the matching degree between the current face and the known person can be accurately judged, improving the accuracy of recognition. The vector library can be continuously updated and expanded to facilitate the addition of new person information and adapt to different application scenarios.
[0158] A500. If the matching similarity is less than the preset threshold, output a face recognition success instruction.
[0159] If the matching similarity is not less than the preset threshold, trigger the second backup strategy to re-collect, or trigger the first backup strategy to upload the face picture and perform recognition based on the uploaded face picture.
[0160] Specifically, a preset threshold can be set, such as 0.8. If the calculated similarity is less than this threshold, it is considered that the face recognition is successful and a face recognition success instruction is output; otherwise, trigger the second collection strategy (such as re-collecting the face image) or the first backup strategy (such as prompting the user to upload a face picture for recognition). By setting the threshold and backup strategies, the situation of recognition failure can be handled, improving the reliability and robustness of the system. Provide a backup plan when recognition fails, giving users the opportunity to re-recognize and enhancing the user experience.
[0161] In this embodiment, through the processing of multiple steps, from face region localization to feature extraction, feature vector generation, and similarity calculation, high-precision face recognition can be achieved to meet the recognition requirements in different scenarios; measures such as fusing continuous frame feature points and adopting backup strategies enable the system to adapt to different image qualities, face postures, and environmental conditions, improving the robustness of the system; the preset vector library can be updated at any time, and the backup strategy can be adjusted according to the actual situation, making the system have good scalability and flexibility and being able to adapt to changing application scenarios.
[0162] Among them, triggering the second backup strategy for re-acquisition specifically includes: prompting the user to adjust the light and angle and then re-recognize, or collecting some target feature points for feature vector acquisition and matching similarity analysis.
[0163] Specifically, some target feature points are processed by a neural network to generate a degraded face feature vector, and the matching similarity of the degraded face feature vector is obtained.
[0164] Further, in this embodiment, when re-acquisition is required, first, a face detection algorithm is used to process the collected image. For example, the Haar cascade classifier can be used, which can identify the face region in the image; the collected color image is converted into a grayscale image, and then the classifier is used to find the face in the grayscale image. If a face is detected, the region where the face is located is extracted. In the detected face region, a feature point extraction algorithm is used to extract the feature points of the face. In practical applications, tools such as the dlib library can be used to extract 68 feature points of the face, and these feature points can accurately describe the shape and contour of the face.
[0165] Under normal circumstances, a trained neural network model is used to process all the extracted feature points to generate a complete face feature vector, which can represent the features of the currently collected face. When the second backup strategy is triggered, some target feature points are collected. For example, the first 30 feature points are selected, and then a special neural network model is used to process these partial feature points to generate a degraded face feature vector. A similarity calculation method, such as cosine similarity, is used to compare the generated feature vector with the vectors in the preset feature vector library to calculate the matching similarity. Cosine similarity can measure the cosine value of the angle between two vectors, and the closer the value is to 1, the higher the similarity.
[0166] Further, this embodiment also includes: obtaining the actual number of re-acquisitions. If the actual number exceeds the preset number, the user is prompted to switch to the manual upload mode.
[0167] In this embodiment, if the similarity does not reach the threshold, it is considered that the face recognition fails. The number of times of re-collection is recorded and incremented by 1, and the user is prompted to continue retrying. The above processes of re-collection and judgment are continuously repeated until the actual number of re-collections exceeds the preset number. When it exceeds the preset number, the user is prompted to switch to the manual upload mode. For example, the message "The number of re-collections exceeds the preset number. Please switch to the manual upload mode" is displayed.
[0168] In this embodiment, determining whether the current browser supports the camera function according to the browser information may include checking whether the target API is available according to the developer tools. If some APIs are not supported, a preset database (Polyfill library) is dynamically loaded for function replacement. If the required API (such as getUserMedia) is not supported and cannot be solved by dynamically loading data, the alternative solution is triggered. The alternative solution includes providing a function that only supports picture upload. (For example, if the camera cannot be used to record videos in real time, a function that only supports picture upload can be provided.) If it is detected that navigator.mediaDevices is not supported, Polyfill can be considered to be used to be compatible with legacy browsers, or a fallback solution (such as only supporting picture upload) can be provided. The target APIs include the APIs for accessing the user's camera and microphone, the APIs for enumerating the currently available multimedia input and output devices, and the APIs representing the constraint conditions supported by the current browser.
[0169] Furthermore, in this application, to meet the needs of different browsers, devices or platforms, this application also includes image processing, specifically including:
[0170] 1) Transfer the image processing logic from the client side to the server side. By building a unified image processing service, it is ensured that all image processing operations are executed in a consistent environment, thereby avoiding the problem of inconsistent processing effects caused by differences in client device performance and browser environment.
[0171] Through this step, the server-side environment is unified, and all image processing operations are executed in the same hardware and software environment, ensuring the consistency of the processing results. High-performance image processing libraries (such as OpenCV, Pillow, etc.) can be deployed on the server side, with stronger processing capabilities and able to execute more complex image processing tasks. The image processing logic is executed on the server side, avoiding potential security risks on the client side, such as code being tampered with or data being leaked.
[0172] 2) Use machine learning algorithms to calibrate and optimize the image processing effect. Specifically, use machine learning algorithms to calibrate and optimize the image processing effects of different devices. The specific method is to collect the image processing results of different devices under different conditions, train a model to predict and adjust the image processing parameters to adapt to the performance differences of different devices, thereby improving the consistency of the image processing effect.
[0173] Through this step, the machine learning model can automatically adjust the image processing parameters according to the characteristics of different devices, has strong adaptability, and can adjust the training objectives of the model according to actual needs, such as optimizing image quality, reducing processing time, etc.; with the continuous accumulation of data, the model can further optimize the image processing effect through continuous learning.
[0174] 3) Perform image processing. Specifically, 3.1) Define a RESTful API interface for receiving the image data and processing parameters uploaded by the client. The interface path is: `POST / api / image / process`; 3.2) Extract the image file and parameters from the request, verify whether the format of the image file (such as whether it is JPEG or PNG) and the size meet the requirements. If they do not meet the requirements, return an error response; 3.3) Decode the uploaded image file from the binary stream into an image object. The specific operation is: use `io.BytesIO` to convert the binary stream into a file object in memory, or use the `Image.open()` method of the Pillow library to read the image data: 3.4) Verify whether the format of the image is supported (such as JPEG, PNG). If the format is not supported, return an error response.
[0175] 4) Extract the cropping parameters (`x`, `y`, `width`, `height`) from the request and verify whether these parameters are valid (such as whether the cropping area exceeds the image range); extract the compression parameters (`target_width`, `target_height`, `quality`) from the request and verify whether these parameters are valid (such as whether the target resolution is reasonable).
[0176] 5) Before outputting the image, check whether its quality meets the requirements (such as resolution, file size, etc.). Specifically, it includes: checking whether the resolution of the processed image meets the target resolution, and checking whether the file size meets the limit (if any); record the key information of the image processing, including input parameters, processing results, and elapsed time.
[0177] 6) Package the processed image data into an HTTP response and return it to the client. If an error occurs during the processing (such as the image format is not supported, the cropping parameters are invalid, etc.), return detailed error information to facilitate the client to perform corresponding processing.
[0178] Furthermore, a method for calibrating and optimizing the image processing effect using machine learning algorithms includes: 1) Collecting image data from different devices (such as mobile phones, tablets, and computers of different models). Ensure that the images cover a variety of shooting conditions, such as different light intensities, background complexities, shooting angles, etc.; for each image, record the hardware information of the device (such as model, camera parameters) and the shooting environment information (such as lighting conditions, background type).
[0179] 2) Annotating the collected images, including the cropped area of the target image, compression quality metrics (such as resolution, file size after compression), and the final image quality score (such as clarity, color fidelity). The annotation can be done manually or assisted by automated tools.
[0180] 3) Using image processing techniques to extract features of the images, including but not limited to: resolution (width and height), color distribution (such as histogram), edge intensity (through edge detection algorithms), clarity (such as calculating the sharpness of the image through the Laplacian operator), file size.
[0181] 4) Extracting device-related features, including: device model, camera parameters (such as pixel, aperture size), operating system version.
[0182] 5) Normalizing the image data, for example, normalizing the pixel values to the range [0, 1]. For example, encoding the device features, such as converting the device model into a numerical feature, and using the annotated target parameters (such as cropped area, compression quality) as labels.
[0183] 6) Selecting a suitable machine learning model according to the problem requirements. For the prediction of image processing parameters, the following models can be considered: traditional machine learning models such as Random Forest, Support Vector Machine (SVM), or Decision Tree; deep learning models such as Convolutional Neural Network (CNN) for image feature extraction and parameter prediction.
[0184] If a deep learning model is selected, design a CNN architecture with the input being image features and device features, and the output being cropping parameters and compression quality parameters; if a traditional machine learning model is selected, design a feature fusion module to combine the image features and device features and then input them into the model.
[0185] 7) Divide the collected data into a training set, a validation set, and a test set, with the ratio being 70%, 15%, 15%; use the training set data to train the model. For deep learning models, use the backpropagation algorithm to optimize the model parameters. For traditional machine learning models, use the cross-validation method to select the optimal hyperparameters.
[0186] 8) Evaluate the performance of the model using the validation set, paying attention to the accuracy and generalization ability metrics. The accuracy is the degree of closeness between the predicted cropping and compression parameters and the labeled values, and the generalization ability is the performance of the model on unseen data. If the model performance is not ideal, adjust the model structure or hyperparameters.
[0187] 9) Deploy the trained model to the server side or the client side. If deployed to the server side, provide an API interface for the client to call; if deployed to the client side, embed the model file into the client application.
[0188] Integrate the model into the image processing pipeline; on the client side or the server side, call the model according to the device characteristics and image characteristics to obtain the cropping and compression parameters, and use the parameters predicted by the model to crop and compress the image.
[0189] Further, a user feedback mechanism can also be designed to allow users to evaluate the quality of the processed image (such as "satisfied" or "not satisfied"), collect user feedback data for subsequent model optimization; regularly collect new data (including user feedback and new device data), retrain the model to optimize performance, and use incremental learning methods to reduce the time and resource consumption of retraining the model.
[0190] Evaluate the performance of the model in actual use, including processing time and image quality metrics (such as the size of the compressed file, image clarity), use the A / B test method to compare the image processing effects before and after optimization. Configure monitoring tools to monitor the running status of the model in real time to ensure its stability and accuracy, and record the logs of model calls, including input features, prediction results, and user feedback, for easy problem troubleshooting and performance optimization.
[0191] Furthermore, for interactions across different platforms, devices, or browsers, they are carried out through DOM operations. Specifically, the dynamic management methods for DOM element states include: 1) Analyze the DOM elements in the page that need to be dynamically managed and their states; common states include showing / hiding loading icons, showing / hiding error messages, updating progress bars, and showing / hiding success messages. 2) Assign a unique identifier (such as a class name or data attribute) to each state for reference in the code. For example, add the class name `loading-icon` to the loading icon and the class name `error-message` to the error message. 3) Encapsulate the DOM operation logic into separate functions, with each function responsible for managing the change of one state. Specifically, create a `showLoading` function to show the loading icon; create a `hideLoading` function to hide the loading icon; create a `showError` function to show the error message; create a `clearError` function to clear the error message. 4) Control the state of the DOM element by adding or removing class names, setting data attributes, or directly manipulating styles. For example, the `showLoading` function shows the loading icon by adding the `visible` class to it, and the `hideLoading` function hides it by removing the `visible` class. 5) When the page loads, initialize the states of all DOM elements that need to be dynamically managed to ensure they are in the default state; specifically, when the page loads, hide all loading icons and error messages. For example, call the `hideLoading` and `clearError` functions in the `DOMContentLoaded` event. 6) Bind listeners to events that need to trigger state changes, such as click events, submit events, or callback functions of asynchronous requests; specifically, when the user clicks the upload button, call the `showLoading` function to show the loading icon. For example, when the `fetch` request is completed, call the `hideLoading` function to hide the loading icon.
[0192] In this embodiment, it also includes calling the corresponding state management function according to the result (success or failure) of the event. If the `fetch` request is successful, call the `clearError` function to clear the error message; if the `fetch` request fails, call the `showError` function to show the error message.
[0193] Furthermore, CSS can be used to enhance the state representation; define the representations of different states in CSS, such as the style when the loading icon is displayed, the style of error prompts, etc.; for example, define the `.loading-icon.visible` style to make the loading icon visible, and define the `.error-message.visible` style to make the error prompt visible.
[0194] Switch the styles of DOM elements by adding or removing class names in JavaScript, thereby achieving state changes; Example: The `showLoading` function shows the loading icon by adding the `visible` class; the `hideLoading` function hides the loading icon by removing the `visible` class.
[0195] Capture possible errors in the state management function and record the error information for debugging; Example: Capture errors in the `showError` function and record them to the console; call the `showError` function in the `.catch()` of the `fetch` request.
[0196] Show the error information to the user in a friendly way, avoiding directly displaying technical error information, such as mapping the error information to a prompt that the user can understand, such as "Network connection failed, please try again later".
[0197] By encapsulating the DOM operation logic into independent functions, reduce code redundancy and improve maintainability; control the state of DOM elements by adding or removing class names, avoiding direct style manipulation; initialize the states of all DOM elements when the page loads to ensure state consistency when the page loads; call the state management function according to the result of the event to ensure that the state change is consistent with the result of the user operation or asynchronous request; define the representations of different states through CSS to make the state change more intuitive and beautiful; capture and handle possible errors, provide friendly error prompts, and avoid user confusion.
[0198] The face recognition method disclosed in this application effectively solves the problems of compatibility, accuracy, performance, and cross-platform generality that exist in implementing face recognition under different platforms through a series of steps. The following is a detailed analysis: When starting face recognition, the solution will obtain the user device information, browser information, and environmental status, which provides comprehensive data support for subsequent judgments and configurations; Different browsers have different degrees of support for Web APIs (such as camera calls), and there are also differences in the cameras and processing capabilities of different devices. By collecting this information, the specific situation of the current usage scenario can be understood in advance.
[0199] The face recognition method disclosed in this application determines whether the current browser supports the camera function based on browser information, and then determines whether there is an available camera in combination with the user device information. This step can effectively avoid invalid operations on browsers that do not support the camera function, or forcibly calling the camera on devices without an available camera, thereby improving the system compatibility; when it is determined that there is no available camera, a first backup strategy is triggered to upload a face picture, and collection is performed based on the uploaded face picture. This method provides an alternative solution for the situation where the camera cannot be used, solves the problem that face recognition cannot be performed due to device problems, and enhances the adaptability of the system.
[0200] After determining that there is an available camera, determine the range of camera parameters supported by the browser according to browser information, including the video stream width range, video stream height range, supported frame rate range, and camera direction range. Different browsers support different camera parameters. Defining these ranges can avoid compatibility problems caused by parameter mismatches; determine the best camera configuration according to the available camera, the range of camera parameters supported by the browser, and the environmental status. Considering the environmental status allows the system to adjust the camera parameters according to the actual situation. For example, appropriately increasing the frame rate or adjusting the resolution in a dim environment can improve the recognition accuracy and performance.
[0201] This solution does not rely on specific platform specifications, but operates based on general browser and device information. Whether in an H5 Web application or on other platforms that support browser calls, this solution can be used for face recognition. This solves the problem that WeChat mini-programs can only run on the WeChat platform and do not have cross-platform universality.
[0202] The solution determines the best camera configuration by obtaining and analyzing various information. Developers can adjust and optimize according to actual needs without having to follow the strict specifications of a specific platform. This provides developers with greater flexibility and freedom, and can better meet the requirements of different application scenarios.
[0203] In summary, the face recognition method of this application effectively solves the problems existing in the prior art when implementing face recognition on different platforms through comprehensive information collection, flexible backup strategies, optimized camera parameter configuration, and good cross-platform universality, and can efficiently and stably implement face recognition on different platforms.
[0204] In a second aspect, this application discloses a face recognition system for executing the face recognition method disclosed in the first aspect of this application, specifically including:
[0205] A response module for obtaining user device information, browser information, and environmental status in response to a face recognition start instruction;
[0206] A judgment module, configured to determine whether the current browser supports the camera function according to the browser information. If so, determine whether there is an available camera according to the user device information. If there is no available camera, trigger a first backup strategy to upload a face picture and perform acquisition based on the uploaded face picture.
[0207] An analysis module, configured to, if there is an available camera, obtain the available camera and determine the range of camera parameters supported by the browser according to the browser information. The range of camera parameters includes the range of video stream width, the range of video stream height, the range of supported frame rates, and the range of camera directions.
[0208] A calling module, configured to determine the best camera configuration according to the available camera, the range of camera parameters supported by the browser, and the environmental status.
[0209] An execution module, configured to call the best camera configuration to obtain face image information and process the face image information based on a preset image analysis strategy to obtain a face recognition result.
[0210] The computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0211] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the face recognition method in the foregoing embodiments of the present disclosure.
[0212] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.
[0213] As Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiments of the present disclosure. Figure 7The computer device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0214] As Figure 7 shown, the computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0215] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 7 a computer device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0216] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from a storage device, or installed from the ROM. When the computer program is executed by the processor, all or part of the steps of the face recognition method of the embodiments of the present disclosure are executed.
[0217] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0218] A computer-readable storage medium according to the embodiments of the present disclosure stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the face recognition methods of the foregoing embodiments of the present disclosure are executed.
[0219] The above computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or removable hard disks), media with built-in rewritable non-volatile memories (e.g., memory cards), and media with built-in ROMs (e.g., ROM cartridges).
[0220] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, which will not be elaborated herein.
[0221] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes, rather than limitations. The above details do not limit the present disclosure to necessarily implement using the above specific details.
[0222] In the present disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0223] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing. So, for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.
[0224] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0225] Various changes, substitutions, and alterations to the technology described herein can be made without departing from the teachings defined by the appended claims. Additionally, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that are currently available or later to be developed that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0226] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0227] The foregoing description has been presented for purposes of illustration and description. Additionally, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although numerous example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A face recognition method, characterized in that: include: In response to the face recognition start instruction, obtain user device information, browser information and environment status; Determine whether the current browser supports the camera function according to the browser information, and if so, determine whether there is an available camera according to the user device information, and if there is no available camera, trigger the first backup strategy to upload the face picture, and collect the face based on the uploaded face picture; If there is an available camera, obtain the available camera, and determine the camera parameter range supported by the browser according to the browser information; The camera parameter range includes a video stream width range, a video stream height range, a supported frame rate range, and a camera direction range; Determine an optimal camera configuration according to the available cameras, the camera parameter range supported by the browser, and the environmental status; The optimal camera configuration is called to obtain facial image information, and the facial image information is processed based on a preset image analysis strategy to obtain a facial recognition result.
2. The face recognition method according to claim 1, characterized in that: If the current browser does not support the camera function, trigger the missing filling strategy, and perform filling library loading processing on the current browser based on the missing filling strategy, so that the current browser supports the camera function after the processing; The triggering of the missing filling strategy and performing filling library loading processing on the current browser based on the missing filling strategy, wherein the processed current browser supports the camera function, includes: Dynamically create a script element, which is used to subsequently load a fill-in library that can fill in the missing functions of the current browser camera; Setting a source address attribute of the script element, wherein the source address points to a network link of a filling library for filling the missing function of the current browser camera; A loading completion event handler is added to the script element. After the script element is successfully loaded, the current browser supports the camera function after the processing.
3. The face recognition method according to claim 1, characterized in that: The triggering of the first backup strategy to upload the face picture includes: Provide a file selector to allow users to select image files and limit the file type to image formats; After the user selects a file, check whether the file size meets the preset requirements. If it exceeds the limit, prompt the user to select a smaller file; After the user selects a picture, display a preview of the picture for the user to confirm; After the user confirms, the uploaded image is preprocessed, and the preprocessing includes one or more of compression and cropping; Create a FormData object, and encapsulate the pre-processed image and necessary form data in the FormData object; the necessary form data includes one or more of a user ID, a request identifier, and an image processing option; Set the request header corresponding to the FormData object; A fetch request is constructed, and a POST request is initiated using the fetch method. The POST request is used to send the FormData object as a request body to a backend API.
4. The face recognition method according to claim 3, characterized in that: If the current browser does not support the camera function, the first backup strategy is triggered to perform face image recognition.
5. The face recognition method according to claim 1, characterized in that: The obtaining of available cameras includes: Get all cameras and their unique identifiers; Perform permission tests on each camera to obtain available cameras.
6. The face recognition method according to claim 3, characterized in that: The processing of the facial image information based on a preset image analysis strategy to obtain a facial recognition result includes: Locate the target face area based on a preset detection strategy; Extracting original feature points in the target face area of each frame, and fusing the original feature points of consecutive frames to obtain target feature points; Processing the target feature points through a neural network to generate a facial feature vector; According to a preset vector library, obtaining the matching similarity of the facial feature vector; If the matching similarity is less than a preset threshold, output a face recognition success instruction; If the matching similarity is not less than a preset threshold, the second backup strategy is triggered to re-collect, or the first backup strategy is triggered to upload a face picture and perform recognition based on the uploaded face picture.
7. The face recognition method according to claim 6, characterized in that: The method of locating the target face area based on a preset detection strategy includes: According to the optimal camera configuration, a fast scan is performed on a low-resolution frame. If no face is detected, the high-resolution frame is switched to perform an accurate scan to obtain a target face. When multiple faces are detected, the face closest to the camera center is selected as the target face; Based on the target face, determining a target boundary, wherein the target boundary does not contain redundant background information; The target face and the target boundary constitute the target face area.
8. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the face recognition method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the face recognition method described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.