In my current project, I needed to fetch data from an API, sanitize it, and return it in a form that is easy to work with. Over the years, I have accumulated snippets and go-to methods for various tasks, but I haven't spent much time refining them. I write code with a focus on performance, readability, and potential enhancements. When I wrote the first version of this code, which was essentially a bunch of if statements, it struck me that for a sequence of conditions, I could use switch(true). I refactored the code, first using the ?? NULL coalescing operator.
Then, I realized that starting with empty() would be better for readability. I should mention that an important principle in my code is "return early," and using empty() at the beginning is ideal.
After trying a few variations, I settled on the following steps:
Here is my refactored function:
function get_ajax_data(){
switch(true){
case !has_valid_nonce():
case empty($_POST['order_id']):
case empty($_POST['ebook_id']):
case empty($_POST['user_id']):
case empty($_POST['action']):
case empty($_POST['format']):
case empty($_POST['order_id']):
case !is_numeric($_POST['order_id']):
case !is_numeric($_POST['ebook_id']):
case !is_numeric($_POST['user_id']):
case !($_POST['format'] === 'epub' || $_POST['format'] === 'mobi'):
return false;
}
return [
'action' => ltrim($_POST['action'],'ebook-insert-client-'),
'format' => $_POST['format'],
'order_id' => $_POST['order_id'],
'ebook_id' => $_POST['ebook_id'],
'user_id' => $_POST['user_id'],
];
}
function has_valid_nonce(){
if( empty($_POST['nonce']) ) return false;
require_once ABSPATH.'wp-includes/pluggable.php';
return wp_verify_nonce( $_POST['nonce'], 'ebook-insert-client-ajax' );
}
The next step was to validate the URL. In PHP, this is pretty straightforward, right? Well, sort of.
case !filter_var( $post['ebook-url'], FILTER_VALIDATE_URL):
Since I’m writing code for WordPress, I could use the wp_http_validate_url() function, but after looking at the source code, I found it to be overly complex and unnecessary for my case. So, I stuck with filter_var().
In this case, I need to receive a URL with a query string with a timestamp. Initially, I used the native PHP function filter_var(), but when I examined what this function considers a valid URL, I realized it was too broad for my use case. So, I moved it into a function and added disqualifying logic:
function is_url_valid($url){
switch( true ){
case $url[0] !== 'h': //to disqualify ftp://
case str_contains($url,'@'): //to early disqualify email
case str_contains($url,'user:'): //to early disqualify login
case str_contains($url,'..'): // 'http://example..com' will pass as a valid URL by var_filter() WTF?!
return false;
default:
return filter_var($url, FILTER_VALIDATE_URL);
}
}
Then I needed to handle parsed URLs. In this scenario, the URL must always have a .epub extension and a numerical value as the query string.
case $parsed_url = parse_url( $post['ebook-url'] ): //early return if URL could not be parsed
case empty($parsed_url['host']):
case empty($parsed_url['query']):
case !is_numeric($parsed_url['query']):
case empty($parsed_url['path']):
case str_contains($parsed_url['path'],'//'): // two slashes "//" will pass as valid URL via filter_var() WTF?!
case !str_ends_with($parsed_url['path'], '.epub'): // in this case only .epub file extension is expected
Yes, I know this might be controversial. Declaring variables in an if statement is generally not considered good practice. However, since this is a switch statement, it looks a bit different and I would argue it's still easy to read and comprehend. If you're using a modern code editor like VS Code, you won't have any issues seeing where the variable comes from, especially given the self-explanatory name.
Without delving into too many details, I found that the slowest part of this process is the WordPress core function wp_verify_nonce(), which takes about 900 microseconds, and I can't do anything about it. Without nonce verification, the rest of the code runs in just 3 microseconds on my local environment. That's pretty efficient, isn't it?
From now on, I plan to use this approach to handle input data, disqualifying malformed or corrupted inputs early and letting only sanitized data pass through. While this article focuses on using switch(true) to quickly filter out bad data, the concept applies to other scenarios too. For instance, when handling inputs like emails or names, I would employ different sanitization techniques to guard against SQL injection and other security risks. The key takeaway is the strategy of catching potential issues early, whether due to incorrect data or malicious attempts.
I'm looking forward to your feedback, are there any weaknesses or areas for improvement, or should I reconsider this approach because there's something better out there? Let me know in the comments below.