Skip to content

Main features

The main features of the das-Peak voice biometrics engine are:

  • Minimum duration: Allows verifying audios with a minimum voice duration of 3 seconds, an industry-record duration. Verification time: 0.14 seconds for comparing a biometric vector and an audio.
  • Text-independent: Allows to compare phrases with different content. That is, the user does not have to remember any phrase or have to read the same phrase to be authenticated.
  • Language-independent: Allows to compare voices in different languages. The biometric voice engine has been specially trained for the following languages: English, Spanish, German, French, Italian, Chinese, Taiwan, Dutch, Estonian, Persian, Turkish, Welsh, Kabyle, Catalan and Euskera.
  • Certified technology: Biometric engine performance rated Internationally by NIST and by the SdSV as one of the leading solutions in the market. (e.g. In the Short-duration Speaker Verification challenge 2020 (SdSV), Veridas achieved a 0.0769 %, being the 2th in single system)
  • Minimum size biometric vector: The biometric vector size is 1.1Kbytes.
  • Voice activity detection: Compute the total quantity of voice in the input audio to accept a verification request.
  • Noise detection: Compute the total quantity of noise in the input audio to accept a verification request.
  • Voice Authenticity Detection: given an audio input it is possible to detect if the voice recorded in the audio is authentic or is a replay attack spoof performed by reproducing a voice through a smartphone speakers or a high fidelity speakers.
  • Different Calibration modes: Include different calibration to obtain better verification/identification results depending on the use case (Telephone channel, Lossless audio, without calibration). If the use case is the telephone audio, the customer must use the parameter calibration="telephone-channel". If the customer uses SDKs for audio capture, the parameter calibration="lossless-audio" should be added.

To ensure optimal performance, recommendation for audios to be captured under specific conditions is needed. VERIDAS offers different proprietary SDKs for audio recording, and they are available for different platforms (iOS, Android). These SDKs ensure that the capture process is performed following the best conditions. Most relevant conditions are:

  • Audio format: WAV.
  • Number of channels: Mono.
  • Bits per Sample: 16.
  • Sampling frequency: 8 kHz or 16 kHz.
  • Do not use lossy audio encoding: audio encodings such as mp3 can degrade the performance of the voice biometry engine. Converting mp3 to wav is not recommended because the lossy encoding is retained.
  • Do not use audios with multiple voices to register/verify: If there are multiple voices in the audio, the voice credential created will contain multiple voice information and this will negatively affect future verification processes.
  • Use "lossless-audio" calibration: To obtain correct verification/identification results working with SDKs, it is necessary to use this calibration.

das-Peak can process any audio that meets the above conditions, not just audios recorded with the SDKs.

System accuracy report

The accuracy of the das-Peak system is given in the document das-Peak Performance Report.